Keep scrapers alive after the first week· Lesson 4 of 6

Bot walls, and what actually changes the outcome

What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.

  • 11 min read
  • No account needed

A protection layer is not trying to work out whether you are a robot. It is scoring how much you cost and how much you are worth, and the score is mostly made of three things: where you are connecting from, how fast you are asking, and whether your request looks like it came from the software it claims to come from.

That framing matters, because it tells you which of your options are real. You cannot argue with the score. You can change the inputs to it, and only some of those inputs are yours to change.

What gets measured, and what you can do about it

SignalWhat it seesYour realistic lever
IP reputationDatacenter range, known provider, prior abuse from that blockRoute differently, or accept the refusal
Request rateRequests per minute from one source, and the regularity of the gapsSlow down and vary. This is the biggest free win.
Header coherenceWhether your headers describe a real client consistentlySend a coherent, honest set. Do not impersonate.
Session behaviourStraight to product pages, no cookies, no referrer, never a listingBehave like a reader: accept cookies, follow links
Browser fingerprintJavaScript challenges, canvas, timingRender properly, or stop

The boring answers outperform the clever ones

Nearly everything written about this topic is about defeating detection, and nearly all of it ages badly, because it describes a specific countermeasure against a specific version of a specific vendor's product. The things that keep working are unglamorous.

Ask less often. Most scraping is dramatically over-frequent relative to how often the underlying data changes. A price that moves twice a week does not need hourly polling, and halving your request rate removes the signal you were most visibly failing on.

Be identifiable. A user agent that says who you are and a contact address is not naivety — it is the thing that gets you an allowlist instead of a ban when somebody looks at the logs. Pretending to be Chrome and failing the fingerprint check is worse than both.

Ask for less. Pulling a thousand product pages when the category listing already carries the price and the title is a thousand requests you did not need, and a hundredfold increase in your visibility.

Cache aggressively. The cheapest request is the one you did not make.

One measured thing about market and locale

Related, and it has cost us real money. Large international retailers decide which market you are in, and that decision changes the price you are served — not the formatting, the actual number.

On Amazon we measured this directly. Sending an Accept-Language header alone did nothing at all: three attempts within the same minute, same market, same price every time. The cookie is what flips the market. Get that wrong and your scraper does not fail — it succeeds, and returns a correct price from the wrong country, which is the silent-partial failure from lesson one wearing a different coat. A single-currency sanity check on the output would have caught it in a day. We did not have one, and the wrong number was live for longer than it should have been.

When to stop

Some sites will not be read, and continuing is a cost with no return. Stop when any two of these are true.

  • You have spent more than a day on one site and the success rate is still under half
  • The fix involves solving a visual or interactive challenge
  • You are being asked to log in, and the terms attached to that login forbid automated access
  • The per-page cost of getting through exceeds what the data is worth to you
  • The thing you need is published somewhere else — a feed, a marketplace listing, an aggregator — that is not fighting you

Worked example: the escalation ladder, priced

Each rung costs more than the one above it, and the order matters because the cheap rungs resolve most cases. Cost is given as a multiple of a plain fetch rather than in currency, because the absolute figures differ by provider and by month while the ratios are the part that generalises — and the ratios are what should govern the decision.

RungWhat it changesRelative costWhen it is the right answer
Slow downRequest rate×1Almost always first — rate is the cheapest thing you control
Fix the headersHeader coherence×1A client whose headers match no real browser
Render the pageExecution of page scriptsabout ×5The value is absent from the source and present on screen
Residential addressIP reputationabout ×10A data centre range is being refused outright
Both togetherBothabout ×25Rarely, and never without a spend cap in front of it
StopNothing×0A documented coverage gap — the next lesson is about writing it

What usually goes wrong

Almost every expensive mistake here is an escalation made before the cheap rungs were tried, or measured somewhere the job does not run.

  • Starting at the expensive rung. Residential addresses and full rendering fix a minority of cases and cost five to twenty-five times as much on every page, permanently.
  • Reproducing it on a laptop. A flagged data centre address is refused outright rather than challenged, so a body fingerprint measured from your machine is a false negative on the host that actually runs the job. Key the detection on the HTTP status.
  • Escalating without a spend cap. The combined rung is roughly twenty-five times the base rate, and a retry loop on top of it is how a month's budget disappears in an afternoon.
  • Changing the market without meaning to. On at least one large retailer the cookie decides the market and the language header alone does nothing, so an escalation can return a perfectly correct price from the wrong country.
  • Impersonating a browser as a strategy. It is an arms race against a team whose whole job it is, and it converts a conduct question into a bad-faith one — course six covers what that costs.
  • Treating stopping as a failure. A documented coverage gap with numbers attached is a result; a partial feed nobody labelled is a liability.

Stopping is a result

A coverage note saying "this retailer is not readable, here is the evidence, here is what we propose instead" is a legitimate deliverable. It is worth considerably more than a feed that is quietly thirty percent wrong because somebody refused to be beaten.

The failure mode to avoid is the one where a site that mostly blocks you produces a trickle of rows and nobody labels them. A partial feed presented as a complete one is the most expensive artefact in this whole discipline.

Next: the failure that does not announce itself — monitoring output rather than exit codes.

Lesson 5: monitor the feed, not the run

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.