Collection, Politeness & Archive

Collection is the boring but most important part: on it depends whether the news will be complete and whether the watching will not turn into a punishment for other people’s sites. Everything here is built around three rules: walk the sources rarely and politely, save what was seen whole, and never multiply one and the same story.

One careful collector

All sources go out to the internet through one shared point. This is deliberate: without it, two dozen sources easily turn into a hundred simultaneous requests, and such crawling quickly leads to being blocked.

  • At most one request per minute to one site — by default. The limit is shared by all the sources looking at that site.
  • Retries with a growing pause — when a site does not answer or answers with an error, the agent waits and tries again instead of hammering on the door every second.
  • Its own proxy and its own headers — can be set for an individual source, when a site demands special conditions.

A repeat request only on change

When a site can answer “nothing changed”, the agent uses it: the page is simply not downloaded again, and the previous result is taken. For you this means less traffic, less load on other people’s sites and fewer reasons to be blocked — with the same result.

The raw-page archive

Every downloaded page is saved whole — as it looked at that moment, with the address and the time. Why: in a month nobody remembers what exactly was written on a competitor’s page, while “prove it” is the most frequent question to news watching.

  • An event card shows where the fact came from — there is a link to the archived copy of exactly that version of the page.
  • Identical texts do not multiply — files are stored by their content, so two identical pages take up space once, not twice.

Repeats and merging

The same story usually arrives from several places: from a feed, from search, from a Telegram channel. Left alone, you would get five cards about the same thing. So the agent looks for repeats by three signs, from the most exact to the softest:

  1. The item’s external id — when the source itself reports a permanent record number, confusing it with another is impossible.
  2. The page address — the same material at the same address.
  3. A text match within 72 hours — reprints and re-publications of the same story within three days.

What comes out of it: repeat publications do not create second records, and the record keeps all the source addresses. That is, the reprint can be opened, but it takes no separate place in the feed.

Full-text completion

Feeds have a known ailment: instead of the article they carry a two-paragraph teaser. The agent notices this and reaches for the article’s full text at its own address — by default, when the teaser is under 500 characters. When the page is locked behind a password or a subscription, what was there stays: a short teaser, without an error and without inventions.

Receiving newsletter material

The product takes no letters itself — there is no mail receiver in it. The scheme is this: in your mail service’s settings you forward the newsletter’s letters to a special intake address the product gives you. The letters arrive signed, and the agent verifies the signature every time — without a valid signature the material is not accepted.

Recovering vanished pages

Sometimes a page disappears: a competitor fixed the price and deleted the old version. When you need exactly that former one, the agent can look for it in an external web archive and put the found copy into the evidence. When there is no copy, it says so instead of substituting random text.

Your own code at the collection step

When the built-in parsing cannot cope, any collection step can be adjusted with your own Python code. Usually this is needed in three cases:

  • Clean a page right after downloading — remove banners, ads and service inserts so they do not look like changes: the code receives the downloaded text and returns the corrected one.
  • Parse a letter in an unusual format — the standard parser is built for ordinary letters; when your newsletter is built differently, your code parses the letter.
  • Enrich an item at the entrance — fix the title, detect the language, add tags, while the item has not yet reached the feed.

Every such run lands in the run journal: visible when it ran, what it received as input and what it returned. When the code breaks, the product honestly stops that step with an error instead of continuing on unverified data. Details — in the Custom Scripts section.

What happens on failures

Failures happen to everyone; what matters is that they do not go unnoticed. So:

  • The source honestly changes state — “a streak of failures”, “error”, “paused”. Not “all good” while the site is silent.
  • A request quota running out is loud — when a source’s limit is exhausted, you learn about it instead of being left with silence and the thought that “nothing is happening”.
  • A global pause is visible — when the check schedule is stopped as a whole, that is shown on the Home screen.

Next → Monitors

← Back to the documentation index