Keep scrapers alive after the first week· Lesson 5 of 6

Monitor the feed, not the run

Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.

  • 11 min read
  • No account needed

Everything in this lesson exists because of one fact: the worst scraper failures do not raise errors. They complete, they write rows, the dashboard refreshes, and the numbers are wrong.

Monitoring the run tells you the job finished. Monitoring the feed tells you whether the thing it produced is usable. Only the second one is worth waking up for, and most teams have only built the first.

Six checks, cheapest first

Implement them in this order. The first two cover most of the real-world damage.

CheckCatchesSuggested trigger
Row count versus trailing averagePagination caps, partial blocks, truncated listsDeviation beyond about 20% of the 7-day mean
Null rate per fieldA single selector breaking while the rest holdAny field whose null rate doubles week on week
Value distributionCurrency flips, unit errors, wrong marketMedian moves more than 15% with no known cause
StalenessA cached response being re-served as freshAny row whose value is byte-identical for an implausible stretch
Coverage against the expected URL listSilent drops you never asked aboutAttempted versus returned falls below your agreed floor
Cross-source agreementEverything else, where you have a second sourceTwo sources disagreeing by more than a tolerance

Row count is the single highest-value check

If you only ever build one of these, build this one. Store the count per source per run and compare it to the trailing average. It is a handful of lines and it catches the majority of silent failures, because almost every silent failure shows up first as "less than usual".

The live case from lesson one is exactly this shape. Rows per week went 10,382 then 8,237 then 8,226, while total requests stayed at 11,032, 11,035 and 11,034. Every page was fetched. Every page was billed. A fifth of the output simply stopped arriving, and because the error count was zero and the job was green, nothing anywhere said so. A row-count delta check would have flagged it in week one.

Thresholds, and the alert nobody reads

A monitor that fires every day is not a monitor, it is a background noise generator, and the second week of it is more dangerous than having no monitor at all because now everyone has learned to dismiss that channel.

Three things keep it honest. Compare against a trailing window rather than a fixed number, so the baseline moves with the business. Suppress alerts for known events — a retailer's seasonal catalogue shrink is not a defect. And separate severities: a 20% drop is a ticket, a 90% drop is a page.

The other discipline is to alert on the derivative, not the level. Nobody can say what the correct row count for a category is. Everybody can tell you that it should not have changed by a fifth overnight.

What to record on every run, from day one

Cheap to write, impossible to reconstruct later.

  • URLs attempted and rows returned, as two separate numbers
  • HTTP status counts, broken out — 403s and 404s mean entirely different things
  • Per-field null counts
  • Which fallback rung answered, if you have a chain
  • Duration, which is the cheapest proxy for "something changed"
  • A content hash per row, so a value that never changes becomes visible

Worked example: the same run, two success rates

One site, one night, one set of numbers. The dashboard reported 100% and the honest figure is about 74%, because the two calculations divide by different things. Nothing here is disputed — both rates are computed from the same run.

MeasureCountRate
URLs on the link list2,667the denominator that matters
Pages attempted2,667
Pages fetched without an error1,964
Rows written1,964
Rows carrying a price, over rows written1,964 of 1,964100% — the figure on the dashboard
Rows carrying a price, over URLs attempted1,964 of 2,667about 74% — the figure to alert on

What usually goes wrong

Monitoring that measures the job rather than the data produces a green dashboard above a feed nobody should be using.

  • Computing a success rate over rows returned. Every row returned has a price by definition, so the rate sits near 100% forever and the check is decorative.
  • Alerting on the level rather than the movement. Catalogues grow and shrink, so a fixed floor either fires every week or never fires; a percentage move against the trailing average tracks what is actually happening.
  • Giving every check the same severity. A doubling in the price null rate and one page returning 404 should not open the same ticket, or the ticket stops being read.
  • Monitoring across all sites at once. The one competitor that disappeared is a few per cent of the total, and a few per cent is indistinguishable from noise.
  • Having no hand check. Twenty rows a month compared against the live pages is the one control in the whole system that cannot itself drift.
  • Not recording attempted counts. Without them the honest denominator above cannot be computed at all, and the dashboard is permanently flattering.

The reconciliation nobody does

Once a month, take twenty rows at random and open the pages by hand. Compare what the feed says to what the page says.

It takes half an hour and it is the only check that catches the category of error where everything is internally consistent and externally wrong — the right price from the wrong market, the list price where the promotion should be, yesterday's number served from a cache. Automated checks compare your data against your data. This is the only step that compares it against reality.

Write the result down with the date. Over a year it becomes the only honest answer you have to "how much do we trust this feed", and that question will be asked by someone senior at the worst possible moment.

Next: the decision itself — when to fix, when to work around, and when to stop.

Lesson 6: decide what to do when a site wins

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.