There are two questions in this lesson and the second one matters more. How often should you check? And how will you find out when checking stops working?
Almost every team gets the first right by accident and the second wrong for months.
How often
Match the cadence to the decision, not to what is technically possible. If your prices are reviewed weekly by a human, a daily feed is already giving them more than they can use, and an hourly feed is giving them seven times the bill for the same decision.
Daily, overnight, is right for the large majority of retail. Weekly is right for slow categories and for the long-tail competitors you decided in lesson two to watch rather than track. Intra-day genuinely matters in two situations: marketplace buy-box repricing, where the feedback loop is minutes, and flash-sale-heavy categories where a competitor's promotion starts and ends inside a day.
One detail worth getting right: run at the same time every day, and record the timestamp with the row. Prices change through the day, and a feed that samples at 03:00 on Monday and 14:00 on Tuesday produces price movements that are artefacts of your own schedule.
Four alerts that catch almost everything
None of these are sophisticated. All of them are better than a green tick on a completed job.
| Alert | Trigger | What it usually means |
|---|---|---|
| Row count drop | Today's rows below 90% of the trailing 7-day median | A competitor redesigned, or added a bot wall, or you lost a category |
| Null rate spike | A field that is normally 98% populated falls below 90% | That specific field moved on the page; the rest of the row is fine |
| Implausible price move | A price changes by more than 50% overnight | Usually a decimal separator or a currency flip, occasionally a real clearance |
| Staleness | A tracked product has not returned a price for 3 consecutive runs | The URL is dead, or it is genuinely discontinued — both need action |
Alert per competitor, not per feed
This is the detail that makes the difference between alerts that work and alerts that get muted. If you have eight competitors and one of them breaks completely, your total row count drops by maybe twelve percent — comfortably inside the noise of a whole-feed threshold, and completely invisible.
Compute the same four checks per competitor, and a total failure becomes a hundred percent drop on one segment, which no threshold can miss. The cost is a slightly longer alert configuration; the benefit is catching the exact failure that a feed-level alert is structurally blind to.
Things that will break your run, roughly in order of likelihood
- A competitor redesigns their product template, usually without warning and usually in Q1
- A bot wall appears, often seasonally, often only for datacentre IP ranges — meaning it works from your laptop and fails in production
- A product URL starts redirecting to a category page, which returns a valid 200 and no price
- A market or language switch changes the currency served from the same URL
- Your own in-scope list grows and nobody updated the expected row count, so the baseline is wrong
- A promotion changes the price element entirely for the duration of a sale, then changes back
Worked example: the morning a feed lost a quarter of itself
Alerts only work if someone wrote down what normal looks like first. One competitor over five days, with the row-count threshold set at 10% below the trailing median. No job failed in that week — every one of the five runs exited zero. Thursday's 531 was a category page that had started paginating differently, and without the row-count check the first sign of it would have been a buyer asking why a product had vanished from the sheet.
| Day | Rows returned | Null price rate | Verdict |
|---|---|---|---|
| Monday | 712 | 1.1% | normal |
| Tuesday | 709 | 0.8% | normal |
| Wednesday | 714 | 1.0% | normal |
| Thursday | 531 | 1.2% | alert — 25% below the 712 median |
| Friday | 528 | 1.0% | second day running — now an incident |
What usually goes wrong
Every item here describes a pipeline whose dashboard is entirely green.
- Alerting on the whole feed instead of per competitor. One competitor out of five dropping to zero is a 20% dip in the total — under almost any sensible threshold, and therefore invisible.
- Treating a green job as a good run. Exit code zero means the code finished, not that the page still contained what it used to contain.
- Testing from your laptop. The competitor serves your office address happily and blocks the data centre the job runs in, so the check passes and the run does not.
- Setting thresholds against a fixed number rather than a trailing median. Catalogues grow, so a static floor either fires every week or stops firing altogether.
- Deleting failed runs. A table containing only successes cannot tell you whether a product was out of stock on Thursday or whether Thursday never happened.
- Setting the price-jump threshold too tight. At 5% every promotional weekend is an incident and the alert gets muted; at 60% you still catch the decimal-point errors and currency mix-ups, which are the ones that actually corrupt a reprice.
Keep the history, including the gaps
Store every observation with its date, rather than overwriting a current price column. The overwrite feels tidier and destroys the asset: price history is where you see that a rival's discount is cyclical, that they follow you within two days, that a category has been drifting down for a quarter.
Store the misses too. A row that returned no price on a given day is information — it may be a stock-out, a delisting or a broken scraper — and a table that only contains successes cannot tell you which. A gap you cannot see is a gap you will interpret as stability.
A trustworthy feed in a system nobody opens still changes nothing. Next: getting it where the decision happens.
Lesson 7: exporting to Sheets, BI and your ERP