Keep scrapers alive after the first week· Lesson 6 of 6

Deciding what to do when a site wins

A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.

  • 10 min read
  • No account needed

At some point a source becomes more expensive than it is worth, and the discipline is to notice that on purpose rather than by exhaustion. Teams that handle this badly do not usually make the wrong call — they make no call at all, and spend six months half-maintaining something nobody has decided to keep.

There are three available answers and it is worth being explicit about which one you are choosing.

The three answers

ChoiceWhen it is rightWhat it costs
RepairCause is identified, fix is bounded, source is load-bearingHours now, and the same hours again at the next redesign
Route aroundThe same data exists somewhere less defendedA different source means different coverage — say so
StopCost per page exceeds the value, or the block is categoricalA documented gap, which is a real cost, honestly priced

Put a number on the data before you argue about the effort

Most of these debates go badly because one side is talking about difficulty and the other about importance, and neither has been quantified.

The question that resolves it: if this source disappeared tomorrow, what decision gets made worse? If the answer is "our weekly price position on four hundred SKUs we actively compete on", that is load-bearing and you should spend real effort. If the answer is "a column on a dashboard that two people open", you have your decision and it is not the one anyone was arguing for.

Then cap the effort before starting. A day, a week, whatever is proportionate — decided in advance, because the sunk cost will absolutely argue for one more afternoon and it will say that every afternoon.

The routing-around checklist

Before concluding a source is unavailable, the same data is often published somewhere nobody thought to look.

  1. 1

    The site's own feed

    Sitemaps, product feeds, affiliate exports and RSS are published deliberately and defended far less than the HTML.

  2. 2

    A marketplace listing

    Plenty of retailers that block you directly also sell on a marketplace that does not, with the same prices attached.

  3. 3

    A comparison site or aggregator

    Lower precision and a lag, but a legitimate second-best, provided you label the provenance.

  4. 4

    The mobile application's backend

    Often a clean JSON API. Check the terms before relying on it — this is where the legal question stops being theoretical, and course six covers it.

  5. 5

    Asking

    Genuinely underused. Some retailers will hand you a feed if you explain what you need and why. A no costs you an email.

How to write a coverage gap so it is useful

A good gap report is three sentences and a number. What is missing, what was tried, what it would take, and what it costs to leave it.

"We cannot read this retailer. Nine hundred pages attempted over two days, all refused at the network level before any content was returned; the pattern is categorical rather than rate-related. Getting through would mean a per-page cost roughly four times our current average, with no guarantee of durability. Leaving it means our price position excludes one of eleven competitors in that segment, which matters most on garden furniture where they are the volume leader."

That is an input to a decision. Compare it to "the scraper for this site doesn't work", which is an apology and tells nobody anything.

The habit underneath all of it: publish what you measured, not what you expected. A page that honestly says the storefront did not return readable product data is worth more than one that pads the gap with a plausible number, because the first can be acted on and the second quietly poisons everything downstream of it.

Worked example: pricing the three answers on one site

A competitor that made up 11% of the feed has gone behind a protection layer. Laid out like this the decision takes an hour. Left undocumented it takes a quarter, and the answer arrived at is usually the one that is not in this table.

RepairRoute aroundStop
What it means hereEscalate to rendering plus a residential addressTake the price from the marketplace listing insteadRemove the site and document the gap
Effortabout two days, then ongoinghalf a dayan hour
Running costabout 25× the base page rate, on 11% of the feedbase rate, different sourcenothing
What the data becomesunchangeda third-party seller's price, not the retailer's — label ita named absence
The risk you are takingit breaks again at the next change, now at 25×the substitute gets treated as equivalent by everyone downstreamsomebody presents a market view with a hole in it
When it winsthe site sets the market price on your top linesthe substitute is genuinely comparablethe site was a benchmark rather than a threat

What usually goes wrong

The three answers are all defensible. What costs money is the fourth thing people do instead.

  • Choosing none of them. The default is a scraper half-fixed every few weeks forever, which costs more than any single column in the table.
  • Arguing about effort before valuing the data. The question is which decision gets worse without this site, and if the answer is none, the right-hand column is free.
  • Routing around without labelling the substitute. A marketplace price and a retailer price are different things, and once they share a column nobody can separate them again.
  • Escalating permanently to solve a temporary problem. The expensive rung tends to stay switched on long after the reason for switching it on has gone.
  • Writing the coverage gap as an apology. "It doesn't work" is not a decision input. "This competitor is 11% of the feed, unavailable since 14 March, substitute available at marketplace level" is.
  • Scheduling no review. A gap documented once and never revisited becomes permanent through neglect rather than through a decision.

The maintenance routine worth having

None of this is clever. All of it compounds.

  • Review the row-count deltas weekly. Ten minutes, and it is the whole early warning system.
  • Refresh URL lists monthly — content drift is constant and silent.
  • Hand-check twenty rows against live pages monthly.
  • Keep a one-line log per source of what broke and what fixed it. It is how a new person becomes useful in a week instead of a quarter.
  • Re-run the value question annually. Sources that were load-bearing stop being load-bearing and nobody notices until someone audits the bill.

Collected data is only useful once rows from different sites can be matched to the same product. That is the next course.

Course 5: matching products across sites

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.