Keep scrapers alive after the first week· Lesson 3 of 6

Selectors that survive a redesign

A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.

  • 11 min read
  • No account needed

Two scrapers can read the same price off the same page and one of them will still be working in two years. The difference is not skill at writing selectors. It is what they chose to point at.

Every extraction target has an implicit contract with the site's front-end team, and most of those contracts are imaginary. A generated class name is not a promise. A deeply nested div path is not a promise. Here is what is ranked by how real the promise is.

Extraction targets, most durable first

"Survives" means it kept working across a front-end rewrite, not just a CSS tweak.

TargetDurabilityWhyCatch
JSON-LD / structured dataVery highIt exists to be machine-read and the SEO team defends itNot every site publishes it, and some publish it wrong
data-testid / data-qa attributesHighTheir own test suite breaks if it changesNot public API — can vanish in a tooling migration
Microdata (itemprop)High where presentSame incentive as JSON-LDIncreasingly rare; measured on 60 major retailers, zero used it
Semantic HTML and ARIA rolesMedium-highAccessibility work tends to survive restylingInconsistently applied
Stable id attributesMediumOften tied to backend templatesFramework rewrites take them out
Human-readable class namesLow-mediumMeaningful to a developer, so sometimes preservedNo guarantee whatsoever
Generated class names (css-1x2y3z)Very lowThey are a build artefactChange on a dependency bump, with no visible change to the page
Positional paths (div > div:nth-child(3))LowestThey encode layout, and layout is what changesBreaks when someone adds a banner

The honest ceiling on structured data

It is also worth knowing how far it gets you, because "just use the schema" is advice people give without having measured it. Across thirteen real storefront fixtures we tested, eleven published nothing machine-readable at all on the product page. Of those that did, several published a JSON-LD block whose price disagreed with the price rendered on screen, usually because it was the list price and the page was showing a promotion.

So: always check, often you will be lucky, and never assume the structured value is the one a customer sees. Cross-check one page by hand before you trust a thousand.

Rules that make a selector last

  • Anchor to meaning, not position. Something that describes what the element is will outlive something that describes where it sits.
  • Shorter is stronger. Each additional step in a path is another thing that can move.
  • Never match on a class that looks generated. If it contains a hash, it is a build artefact.
  • Prefer an attribute the site's own tests depend on. Their CI is now protecting your scraper for free.
  • Extract the raw value and convert later. Pull the number and the currency as they appear, and do unit conversion as a separate step — then a formatting change on the site does not become a parsing bug.
  • Write down what you expect. A note saying "price is in the JSON-LD offers block, fallback is the h1 sibling" turns a future outage into a five-minute fix for whoever is on call.

Worked example: a chain that reports its own decay

Four rungs on one site, with the share of pages each one answered in January and again in June. The chain never failed and nothing ever alerted. What changed is that the feed moved off a source somebody defends and onto one nobody does — which is precisely the state the next redesign will break. It is visible only because each run recorded which rung produced the answer.

RungSourceShare in JanuaryShare in June
1JSON-LD Product, offers.price62%0%
2A data-testid attribute31%58%
3A CSS class selector6%39%
4Regex over the visible text1%3%
Pages that produced a price100%100%

What usually goes wrong

A selector is written once and then lives for years, so the failure modes are all about time rather than correctness.

  • Writing a selector against a generated class name. Those strings are build output: nobody is defending them and they change without any person deciding to change them.
  • Assuming structured data will be there. On one measured set of thirteen storefront fixtures, eleven published nothing machine-readable on the product page. Check, rather than designing around the assumption.
  • Trusting structured data that is there without comparing it to the page. It is the most stable source available and also the one most often left stale after a promotion.
  • Building a silent fallback chain. It converts a loud failure into a gradual decay, and you discover it when the last rung goes as well.
  • Taking the first match. A page can carry the same price string in a header, a schema block and a recently-viewed carousel; counting votes across candidates is more robust than taking whichever came first.
  • Selecting on position. Third div inside the second section works perfectly until somebody adds a banner, which is a thing marketing does without telling engineering.

Build a fallback chain, but make it noisy

The strong pattern is an ordered chain: try the structured data, then the test attribute, then the semantic element, then the hand-written selector. First one that yields a plausible value wins.

The pattern has one failure mode and it is severe. A silent chain hides decay. If rung one stopped working in March and rung four has been quietly covering for it ever since, you will find out in September when rung four breaks too — and by then nobody remembers what rung one was for.

So record which rung answered. Not an alert, just a field on the row. When the distribution shifts — eighty percent of rows suddenly answering from rung three instead of rung one — that is the redesign, caught weeks before it would have surfaced as an outage.

One more detail from a bug that reached production: when a chain votes, count the votes rather than taking the first match. A page can contain the same attribute fifty times, with forty-nine of them empty stubs. First-match returns the stub. So does last-match. Counting what the rungs actually agree on returns the price.

Next: what a bot wall is actually measuring, and what honestly changes the outcome.

Lesson 4: bot walls and what actually works

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.