[{"data":1,"prerenderedAt":245},["ShallowReactive",2],{"learn-lesson-keep-scrapers-alive-selectors-that-survive-a-redesign":3},{"course":4,"lesson":67,"index":213,"outline":214,"prev":243,"next":244},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"keep-scrapers-alive",4,"You already have something running","6 lessons, about 65 minutes","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Keep Scrapers Alive: A Free 6-Lesson Maintenance Course","Why web scrapers break and what to do about it. Reading failures properly, selectors that survive redesigns, bot walls, monitoring the feed rather than the run, and knowing when to stop. Free, ungated.","web scraper maintenance, scraper broke, scraper monitoring, css selector best practice, bot detection, scraper error handling, data quality monitoring, web scraping reliability","A free course on why scrapers break and how to keep them running","Six written lessons on scraper maintenance: failure diagnosis, durable selectors, bot walls, feed monitoring, and deciding when a site has won.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course four","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.",{"label":21,"url":22},"Start with lesson one","/learn/keep-scrapers-alive/why-scrapers-break",{"label":24,"url":25},"See real run output","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Name the five reasons a scraper stops working, and tell them apart from the evidence rather than from a hunch","Read a failure down to its cause instead of restarting the job and hoping","Write selectors that survive a front-end rewrite, and recognise the ones that will not","Understand what a bot wall is actually measuring, and what genuinely changes the outcome","Monitor the output of a feed, not just whether the job exited zero","Decide, with a rule rather than a mood, when to stop fighting a site",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Anyone who already has scrapers running and keeps getting surprised by them","Developers who inherited somebody else's collection layer and want to stop firefighting it","Analysts who depend on a scraped feed and need to know how much to trust it","Not written for",[44,45],"First-time scraper authors — start with course one or course three, this assumes something already runs","Anyone looking for techniques to defeat a specific site's protection. That is not what this is, and lesson four explains why that framing loses.",{"title":47,"intro":48},"The six lessons","Read in order the first time. After that it works as a reference — lesson two is the one people come back to.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","What people usually want to know when a scraper has just broken.",[54,57,60,63],{"title":55,"description":56},"My scraper broke this morning. Which lesson do I read?","Lesson two. It is the diagnosis lesson and it is deliberately the longest. Most wasted maintenance time comes from fixing the wrong thing — rewriting a selector when the page never loaded, or rotating proxies when the product was simply discontinued.",{"title":58,"description":59},"Is this about avoiding detection?","No. Lesson four covers bot walls because you will meet them, but the framing there is about what the wall is measuring and what honestly changes the outcome, not about evasion. The most reliable answers turn out to be boring: slow down, be identifiable, and accept that some sites will say no.",{"title":61,"description":62},"Does any of this apply if I use a managed service?","The diagnosis and monitoring lessons especially. A vendor can own the fetching and the extraction, but nobody else can know that your category page should have returned two thousand rows and came back with four hundred. That judgement stays with you whatever you buy.",{"title":64,"description":65},"How much maintenance should I actually expect?","Measured across our own production groups, a storefront with no major redesign needs attention a handful of times a year, and a redesign costs a few hours per affected site. The trap is not the hours. It is that they arrive unannounced, all at once, and usually on the day something else is already on fire.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":204,"next_step":209},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.","11 min",false,{"title":75,"description":76,"keywords":77},"CSS Selectors That Survive a Website Redesign","Which extraction targets are durable and which break on the next deploy. Structured data, test IDs, semantic markup and generated class names ranked, plus how to build a fallback chain.","css selector best practice, xpath vs css selector, scraper selector broke, json-ld scraping, structured data extraction, durable web scraping",[79,84,134,139,144,154,187,197],{"type":80,"paragraphs":81},"prose",[82,83],"Two scrapers can read the same price off the same page and one of them will still be working in two years. The difference is not skill at writing selectors. It is what they chose to point at.","Every extraction target has an implicit contract with the site's front-end team, and most of those contracts are imaginary. A generated class name is not a promise. A deeply nested div path is not a promise. Here is what is ranked by how real the promise is.",{"type":85,"title":86,"intro":87,"headers":88,"rows":93},"table","Extraction targets, most durable first","\"Survives\" means it kept working across a front-end rewrite, not just a CSS tweak.",[89,90,91,92],"Target","Durability","Why","Catch",[94,99,104,109,114,119,124,129],[95,96,97,98],"JSON-LD / structured data","Very high","It exists to be machine-read and the SEO team defends it","Not every site publishes it, and some publish it wrong",[100,101,102,103],"data-testid / data-qa attributes","High","Their own test suite breaks if it changes","Not public API — can vanish in a tooling migration",[105,106,107,108],"Microdata (itemprop)","High where present","Same incentive as JSON-LD","Increasingly rare; measured on 60 major retailers, zero used it",[110,111,112,113],"Semantic HTML and ARIA roles","Medium-high","Accessibility work tends to survive restyling","Inconsistently applied",[115,116,117,118],"Stable id attributes","Medium","Often tied to backend templates","Framework rewrites take them out",[120,121,122,123],"Human-readable class names","Low-medium","Meaningful to a developer, so sometimes preserved","No guarantee whatsoever",[125,126,127,128],"Generated class names (css-1x2y3z)","Very low","They are a build artefact","Change on a dependency bump, with no visible change to the page",[130,131,132,133],"Positional paths (div > div:nth-child(3))","Lowest","They encode layout, and layout is what changes","Breaks when someone adds a banner",{"type":135,"variant":136,"title":137,"text":138},"callout","note","Check for structured data before you write anything","Fetch the page and search the source for application/ld+json. If a Product block is there with a price in it, you are done, and that extraction will likely outlive three redesigns. It takes thirty seconds to check and people routinely skip it and hand-write a selector against markup that will be gone by spring.",{"type":80,"title":140,"paragraphs":141},"The honest ceiling on structured data",[142,143],"It is also worth knowing how far it gets you, because \"just use the schema\" is advice people give without having measured it. Across thirteen real storefront fixtures we tested, eleven published nothing machine-readable at all on the product page. Of those that did, several published a JSON-LD block whose price disagreed with the price rendered on screen, usually because it was the list price and the page was showing a promotion.","So: always check, often you will be lucky, and never assume the structured value is the one a customer sees. Cross-check one page by hand before you trust a thousand.",{"type":145,"title":146,"items":147},"list","Rules that make a selector last",[148,149,150,151,152,153],"Anchor to meaning, not position. Something that describes what the element is will outlive something that describes where it sits.","Shorter is stronger. Each additional step in a path is another thing that can move.","Never match on a class that looks generated. If it contains a hash, it is a build artefact.","Prefer an attribute the site's own tests depend on. Their CI is now protecting your scraper for free.","Extract the raw value and convert later. Pull the number and the currency as they appear, and do unit conversion as a separate step — then a formatting change on the site does not become a parsing bug.","Write down what you expect. A note saying \"price is in the JSON-LD offers block, fallback is the h1 sibling\" turns a future outage into a five-minute fix for whoever is on call.",{"type":85,"title":155,"intro":156,"headers":157,"rows":162},"Worked example: a chain that reports its own decay","Four rungs on one site, with the share of pages each one answered in January and again in June. The chain never failed and nothing ever alerted. What changed is that the feed moved off a source somebody defends and onto one nobody does — which is precisely the state the next redesign will break. It is visible only because each run recorded which rung produced the answer.",[158,159,160,161],"Rung","Source","Share in January","Share in June",[163,168,173,178,183],[164,165,166,167],"1","JSON-LD Product, offers.price","62%","0%",[169,170,171,172],"2","A data-testid attribute","31%","58%",[174,175,176,177],"3","A CSS class selector","6%","39%",[179,180,181,182],"4","Regex over the visible text","1%","3%",[184,185,186,186],"","Pages that produced a price","100%",{"type":145,"title":188,"intro":189,"items":190},"What usually goes wrong","A selector is written once and then lives for years, so the failure modes are all about time rather than correctness.",[191,192,193,194,195,196],"Writing a selector against a generated class name. Those strings are build output: nobody is defending them and they change without any person deciding to change them.","Assuming structured data will be there. On one measured set of thirteen storefront fixtures, eleven published nothing machine-readable on the product page. Check, rather than designing around the assumption.","Trusting structured data that is there without comparing it to the page. It is the most stable source available and also the one most often left stale after a promotion.","Building a silent fallback chain. It converts a loud failure into a gradual decay, and you discover it when the last rung goes as well.","Taking the first match. A page can carry the same price string in a header, a schema block and a recently-viewed carousel; counting votes across candidates is more robust than taking whichever came first.","Selecting on position. Third div inside the second section works perfectly until somebody adds a banner, which is a thing marketing does without telling engineering.",{"type":80,"title":198,"paragraphs":199},"Build a fallback chain, but make it noisy",[200,201,202,203],"The strong pattern is an ordered chain: try the structured data, then the test attribute, then the semantic element, then the hand-written selector. First one that yields a plausible value wins.","The pattern has one failure mode and it is severe. A silent chain hides decay. If rung one stopped working in March and rung four has been quietly covering for it ever since, you will find out in September when rung four breaks too — and by then nobody remembers what rung one was for.","So record which rung answered. Not an alert, just a field on the row. When the distribution shifts — eighty percent of rows suddenly answering from rung three instead of rung one — that is the redesign, caught weeks before it would have surfaced as an outage.","One more detail from a bug that reached production: when a chain votes, count the votes rather than taking the first match. A page can contain the same attribute fifty times, with forty-nine of them empty stubs. First-match returns the stub. So does last-match. Counting what the rungs actually agree on returns the price.",[205,206,207,208],"Rank targets by whose incentive protects them: structured data and test IDs have defenders, generated class names do not.","Check for a JSON-LD Product block before writing any selector — but expect most sites not to have one.","Measured: 11 of 13 storefront fixtures published nothing machine-readable on the product page.","Fallback chains should record which rung answered, or they hide decay until every rung is gone.",{"text":210,"label":211,"url":212},"Next: what a bot wall is actually measuring, and what honestly changes the outcome.","Lesson 4: bot walls and what actually works","/learn/keep-scrapers-alive/bot-walls-and-what-actually-works",2,[215,221,227,228,233,238],{"slug":216,"navTitle":217,"title":218,"summary":219,"time":220,"needsAccount":73},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.","10 min",{"slug":222,"navTitle":223,"title":224,"summary":225,"time":226,"needsAccount":73},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.","12 min",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":229,"navTitle":230,"title":231,"summary":232,"time":72,"needsAccount":73},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":234,"navTitle":235,"title":236,"summary":237,"time":72,"needsAccount":73},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":239,"navTitle":240,"title":241,"summary":242,"time":220,"needsAccount":73},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"slug":222,"navTitle":223,"title":224,"summary":225,"time":226,"needsAccount":73},{"slug":229,"navTitle":230,"title":231,"summary":232,"time":72,"needsAccount":73},1791047867249]