Keep scrapers alive after the first week· Lesson 1 of 6

The five reasons a scraper stops working

Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.

  • 10 min read
  • No account needed

A scraper that worked yesterday and does not work today has failed for one of five reasons. They look similar from the outside — an empty result, or fewer rows than usual — and they have almost nothing in common underneath. The habit worth building is naming the category before touching any code, because four of the five fixes are wrong four fifths of the time.

In rough order of how often they occur on a mature setup:

The five failure modes

Frequency here is from our own production groups, not an industry figure. Your mix will differ by sector; the categories will not.

FailureWhat you seeWhat actually changedTypical fix
Content driftFewer rows, no errorsProducts delisted, renamed or moved; the site is fineNothing. Update the URL list.
RedesignZero rows, page fetched fineThe markup moved; your selector points at nothingRe-point the selector
Bot wallNon-200 status, or a page that is not the product pageThe site decided you are not a browserSlow down, identify yourself, or route differently
Rendering changePage fetched, markup present, values emptyContent moved behind JavaScriptRender the page, or read the data source directly
Silent partialLooks normal, is wrongPagination capped, a region defaulted, a currency flippedThe hardest. Lesson five.

The one that costs the most is not the one you expect

Redesigns feel like the enemy because they are dramatic — everything goes to zero and somebody notices within the hour. That visibility is a gift. A loud failure is a cheap failure.

The expensive category is the last row. A run that returns four hundred rows where it used to return two thousand exits successfully. The scheduler is happy. The dashboard has data in it. And every number downstream is now computed over a fifth of the catalogue, which is worse than having no number at all, because somebody will price against it.

We have measured exactly this on a live client group: rows per week fell from 10,382 to 8,226 across two weeks while the total number of requests stayed identical — 11,032 then 11,034. Nothing errored. Nothing alerted. The job had been fetching and billing for every page the whole time and simply returning less.

Questions that place a failure in the right category in under a minute

Ask these in order. The first one that gives an interesting answer is usually the whole diagnosis.

  • What HTTP status came back? A non-200 is a bot wall or an outage, and no selector work will help.
  • Did the page body arrive, and is it the page you asked for? A 200 that returns a challenge page is still a block.
  • Is the field present in the markup but empty, or absent entirely? Present-but-empty points at rendering; absent points at a redesign.
  • Did every URL fail, or a subset? A subset is almost always content drift.
  • Did the row count fall, or go to zero? Zero is loud and simple. A fall is the dangerous one.

Worked example: one week of symptoms, placed in five minutes

Five scrapers, five symptoms, five different fixes. The column to read is the middle one: the evidence that separates the categories is cheap to collect, and almost nobody collects it, which is why maintenance feels like one endless problem rather than five bounded ones.

SymptomThe evidence that settles itCategory, and what the fix is
Zero rows, job red, connection refusedHTTP status on a single URLBot wall — slow down, or stop and document the gap
Zero rows, job green, 200 returnedIs the price in the HTML source, or only on screen?Rendering change — the page went client-side
Price null on every row, everything else intactSearch the response body for the literal price stringContent drift — one selector, five minutes
Every field null, 200 returned, body unrecognisableOpen the page in a browserRedesign — rewrite the extraction for this site
Rows down 21% over two weeks, no errors anywhereRow count against the trailing average, per siteSilent partial — the expensive one

What usually goes wrong

Four of these five are attempts to fix a scraper before anyone has said which of the five things is wrong with it.

  • Treating all five as one problem called maintenance. They have four different fixes and one of them is "stop", so the category has to be named before anyone touches code.
  • Reaching for proxies first. It is the standard answer to one of the five modes, the wrong answer to the other four, and expensive in all five.
  • Using exit code zero as the health signal. The worst mode in the table above returns successfully, every day, for a fortnight.
  • Diagnosing from a browser. Yours carries cookies, history and a residential address; the job has none of them, which is why the page you are looking at is not the page it received.
  • Rewriting the selector before checking whether the value is in the source at all. If the price only exists after JavaScript runs, no selector will ever find it, and you can spend a day proving that.

Why it matters that these are five and not one

Because the reflex fix for each is useless for the other four. Rotating proxies does nothing about a redesign. Rewriting a selector does nothing about a product that no longer exists. Adding a browser renderer to a site that was never JavaScript-driven costs you money per page and fixes nothing.

Most teams that describe scraping as "constant maintenance" are not doing more maintenance than anyone else. They are doing the same amount, spent on the wrong category, which means the actual cause survives the fix and comes back next week.

Next: how to actually read a failure instead of guessing which of the five it was.

Lesson 2: read the failure, not the symptom

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.