The instinct when a run comes back empty is to open the page in a browser, see the price sitting there in plain sight, and conclude that the selector is broken. That conclusion is right perhaps a third of the time, and the other two thirds are expensive, because you will rewrite a selector that was never wrong and the run will stay broken.
The routine below is ordered so that each step can only be reached if the previous one has been eliminated. It is deliberately mechanical. The whole point is to remove the guessing.
The routine
Do these in order on a single URL. Never debug against the whole run — you cannot see anything in five thousand rows of log.
- 1
Pick one failing URL and work only on that
One URL that definitely used to work. If you cannot name one, that is itself the finding: you may be looking at content drift rather than breakage.
- 2
Check the status code before anything else
A 403, 429 or 503 ends the investigation. That is a bot wall or a rate limit, and nothing in your extraction layer is involved. Jump to lesson four.
- 3
Check what the body actually is
A 200 is not proof you got the product page. Challenge pages, consent interstitials, geo-redirects and soft 404s all return 200 with a perfectly valid HTML body. Look at the title tag and the length. A product page that is suddenly 4 KB is not a product page.
- 4
Search the raw body for the value, not for your selector
Take the price you can see in the browser and grep the fetched body for the digits. If the number is in there, you have an extraction problem. If it is not, the page you fetched is not the page you looked at, and that is a rendering or routing problem.
- 5
Only now look at the selector
You have earned the right to. And at this point the fix is usually obvious, because you know the value is present and you know where.
- 6
Re-run the single URL, then ten, then the group
Fixing one and immediately launching five thousand is how a wrong fix becomes an expensive wrong fix.
Three false conclusions this routine is built to prevent
The first is "it works in my browser, so the site is fine". Your browser has cookies, a residential IP, a full JavaScript engine and a history with that domain. It is not a control group. A page that loads for you and refuses a datacenter address is the ordinary case, not an anomaly — and it means a failure you cannot reproduce locally is still real.
The second is "it returned 200, so we got the page". Covered above, and it is the one that burns the most hours, because a 200 feels conclusive.
The third is subtler: "the fix worked, the run is green". Green after a fix means the run completed. It does not mean the rows are right. Compare the row count to last week's before you call it done — a selector that now matches a different element will happily return five thousand rows of the wrong thing.
Symptom to cause, once you have the evidence
| Evidence | Cause | Where to go |
|---|---|---|
| 403 / 429 / 503 | Bot wall or rate limit | Lesson 4 |
| 200, tiny body, unexpected title | Challenge or consent interstitial | Lesson 4 |
| 200, full body, value absent from source | Rendered client-side | Lesson 3 |
| 200, value present in body, selector misses | Redesign | Lesson 3 |
| Some URLs fine, some 404 | Content drift | Refresh the URL list |
| Everything green, fewer rows | Silent partial failure | Lesson 5 |
Worked example: ten minutes on one URL
A feed that wrote 1,964 rows from 2,667 attempted pages. The temptation is to theorise — delisted products, a bad selector, a blocked range — and every theory sounds plausible. The routine answers it without a theory, and no step costs more than two minutes.
- 1
Check the status distribution across the whole run, not a sample
If the failures all carry one status class you have the answer in ninety seconds. In this case every failure was a page-level error and not one was an empty result, which already kills the delisted-products theory: a delisted page returns something, it does not fail.
- 2
Fetch one failing URL and read the status
A 403 is a bot wall. A 404 is a stale URL list. A 200 means the cause is further down the page and the next two steps are worth the time.
- 3
Search the raw body for the literal value
Before looking at any selector, search the response for the price exactly as it appears on screen. Present means a selector problem. Absent means a rendering problem, and no amount of selector work will ever fix it.
- 4
Only now look at the selector
Three of the four candidate causes have been eliminated for the price of four fetches. This is the point at which editing the selector is an informed act rather than a guess.
- 5
Prove the path ran before explaining why it failed
If a fix was deployed and nothing changed, establish that the new code executed at all — a timing difference, a log line, a billing delta at the provider. Explaining the failure of a code path that never ran is the most expensive way to spend an afternoon that is available to anybody.
What usually goes wrong
Each of these turns a ten-minute diagnosis into a two-day one.
- Theorising before fetching. Delisted products, a supplier feed change and a broken selector all explain the same row count, and a single fetch separates them.
- Reading the status and stopping there. A 200 routinely carries a consent interstitial, a geo-redirect or a challenge page, all of which parse to nothing at all.
- Testing from your own machine and concluding the site is fine. Your laptop has cookies, history and a residential address; the job has a flagged data centre range. The site treats them as different visitors, because they are.
- Sampling the failures. The status distribution across the whole run is one query and it frequently ends the investigation by itself.
- Changing three things and redeploying. When it works nobody knows which change worked, and when it does not nobody knows which one to undo.
- Explaining a failure in code that never executed. Prove it ran — timing, logs, or a billing delta — before explaining anything about it.
Prove the path ran before you explain why it failed
This is the rule that took us longest to learn and it generalises past scraping. Before building any theory about why a step failed, establish that the step executed at all.
Timing is the cheapest evidence: a fetch that supposedly hit a renderer and returned in 120 milliseconds did not hit a renderer. Billing is the next cheapest, if your provider exposes a credit balance — make one known-good call, note the delta, then diff the balance around the request you are investigating. If the counter did not move, the code path you are theorising about never ran, and every theory about its behaviour is noise.
We shipped one wrong fix and one wrong explanation in a single session by skipping that check. Both were plausible. Neither touched the actual cause.
Next: the extraction side — which selectors survive a rewrite and which are built to fail.
Lesson 3: selectors that survive a redesign