Everything in this lesson exists because of one fact: the worst scraper failures do not raise errors. They complete, they write rows, the dashboard refreshes, and the numbers are wrong.
Monitoring the run tells you the job finished. Monitoring the feed tells you whether the thing it produced is usable. Only the second one is worth waking up for, and most teams have only built the first.
Six checks, cheapest first
Implement them in this order. The first two cover most of the real-world damage.
| Check | Catches | Suggested trigger |
|---|---|---|
| Row count versus trailing average | Pagination caps, partial blocks, truncated lists | Deviation beyond about 20% of the 7-day mean |
| Null rate per field | A single selector breaking while the rest hold | Any field whose null rate doubles week on week |
| Value distribution | Currency flips, unit errors, wrong market | Median moves more than 15% with no known cause |
| Staleness | A cached response being re-served as fresh | Any row whose value is byte-identical for an implausible stretch |
| Coverage against the expected URL list | Silent drops you never asked about | Attempted versus returned falls below your agreed floor |
| Cross-source agreement | Everything else, where you have a second source | Two sources disagreeing by more than a tolerance |
Row count is the single highest-value check
If you only ever build one of these, build this one. Store the count per source per run and compare it to the trailing average. It is a handful of lines and it catches the majority of silent failures, because almost every silent failure shows up first as "less than usual".
The live case from lesson one is exactly this shape. Rows per week went 10,382 then 8,237 then 8,226, while total requests stayed at 11,032, 11,035 and 11,034. Every page was fetched. Every page was billed. A fifth of the output simply stopped arriving, and because the error count was zero and the job was green, nothing anywhere said so. A row-count delta check would have flagged it in week one.
Thresholds, and the alert nobody reads
A monitor that fires every day is not a monitor, it is a background noise generator, and the second week of it is more dangerous than having no monitor at all because now everyone has learned to dismiss that channel.
Three things keep it honest. Compare against a trailing window rather than a fixed number, so the baseline moves with the business. Suppress alerts for known events — a retailer's seasonal catalogue shrink is not a defect. And separate severities: a 20% drop is a ticket, a 90% drop is a page.
The other discipline is to alert on the derivative, not the level. Nobody can say what the correct row count for a category is. Everybody can tell you that it should not have changed by a fifth overnight.
What to record on every run, from day one
Cheap to write, impossible to reconstruct later.
- URLs attempted and rows returned, as two separate numbers
- HTTP status counts, broken out — 403s and 404s mean entirely different things
- Per-field null counts
- Which fallback rung answered, if you have a chain
- Duration, which is the cheapest proxy for "something changed"
- A content hash per row, so a value that never changes becomes visible
Worked example: the same run, two success rates
One site, one night, one set of numbers. The dashboard reported 100% and the honest figure is about 74%, because the two calculations divide by different things. Nothing here is disputed — both rates are computed from the same run.
| Measure | Count | Rate |
|---|---|---|
| URLs on the link list | 2,667 | the denominator that matters |
| Pages attempted | 2,667 | |
| Pages fetched without an error | 1,964 | |
| Rows written | 1,964 | |
| Rows carrying a price, over rows written | 1,964 of 1,964 | 100% — the figure on the dashboard |
| Rows carrying a price, over URLs attempted | 1,964 of 2,667 | about 74% — the figure to alert on |
What usually goes wrong
Monitoring that measures the job rather than the data produces a green dashboard above a feed nobody should be using.
- Computing a success rate over rows returned. Every row returned has a price by definition, so the rate sits near 100% forever and the check is decorative.
- Alerting on the level rather than the movement. Catalogues grow and shrink, so a fixed floor either fires every week or never fires; a percentage move against the trailing average tracks what is actually happening.
- Giving every check the same severity. A doubling in the price null rate and one page returning 404 should not open the same ticket, or the ticket stops being read.
- Monitoring across all sites at once. The one competitor that disappeared is a few per cent of the total, and a few per cent is indistinguishable from noise.
- Having no hand check. Twenty rows a month compared against the live pages is the one control in the whole system that cannot itself drift.
- Not recording attempted counts. Without them the honest denominator above cannot be computed at all, and the dashboard is permanently flattering.
The reconciliation nobody does
Once a month, take twenty rows at random and open the pages by hand. Compare what the feed says to what the page says.
It takes half an hour and it is the only check that catches the category of error where everything is internally consistent and externally wrong — the right price from the wrong market, the list price where the promotion should be, yesterday's number served from a cache. Automated checks compare your data against your data. This is the only step that compares it against reality.
Write the result down with the date. Over a year it becomes the only honest answer you have to "how much do we trust this feed", and that question will be asked by someone senior at the worst possible moment.
Next: the decision itself — when to fix, when to work around, and when to stop.
Lesson 6: decide what to do when a site wins