[{"data":1,"prerenderedAt":224},["ShallowReactive",2],{"learn-lesson-keep-scrapers-alive-read-the-failure-not-the-symptom":3},{"course":4,"lesson":67,"index":192,"outline":193,"prev":222,"next":223},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"keep-scrapers-alive",4,"You already have something running","6 lessons, about 65 minutes","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Keep Scrapers Alive: A Free 6-Lesson Maintenance Course","Why web scrapers break and what to do about it. Reading failures properly, selectors that survive redesigns, bot walls, monitoring the feed rather than the run, and knowing when to stop. Free, ungated.","web scraper maintenance, scraper broke, scraper monitoring, css selector best practice, bot detection, scraper error handling, data quality monitoring, web scraping reliability","A free course on why scrapers break and how to keep them running","Six written lessons on scraper maintenance: failure diagnosis, durable selectors, bot walls, feed monitoring, and deciding when a site has won.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course four","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.",{"label":21,"url":22},"Start with lesson one","/learn/keep-scrapers-alive/why-scrapers-break",{"label":24,"url":25},"See real run output","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Name the five reasons a scraper stops working, and tell them apart from the evidence rather than from a hunch","Read a failure down to its cause instead of restarting the job and hoping","Write selectors that survive a front-end rewrite, and recognise the ones that will not","Understand what a bot wall is actually measuring, and what genuinely changes the outcome","Monitor the output of a feed, not just whether the job exited zero","Decide, with a rule rather than a mood, when to stop fighting a site",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Anyone who already has scrapers running and keeps getting surprised by them","Developers who inherited somebody else's collection layer and want to stop firefighting it","Analysts who depend on a scraped feed and need to know how much to trust it","Not written for",[44,45],"First-time scraper authors — start with course one or course three, this assumes something already runs","Anyone looking for techniques to defeat a specific site's protection. That is not what this is, and lesson four explains why that framing loses.",{"title":47,"intro":48},"The six lessons","Read in order the first time. After that it works as a reference — lesson two is the one people come back to.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","What people usually want to know when a scraper has just broken.",[54,57,60,63],{"title":55,"description":56},"My scraper broke this morning. Which lesson do I read?","Lesson two. It is the diagnosis lesson and it is deliberately the longest. Most wasted maintenance time comes from fixing the wrong thing — rewriting a selector when the page never loaded, or rotating proxies when the product was simply discontinued.",{"title":58,"description":59},"Is this about avoiding detection?","No. Lesson four covers bot walls because you will meet them, but the framing there is about what the wall is measuring and what honestly changes the outcome, not about evasion. The most reliable answers turn out to be boring: slow down, be identifiable, and accept that some sites will say no.",{"title":61,"description":62},"Does any of this apply if I use a managed service?","The diagnosis and monitoring lessons especially. A vendor can own the fetching and the extraction, but nobody else can know that your category page should have returned two thousand rows and came back with four hundred. That judgement stays with you whatever you buy.",{"title":64,"description":65},"How much maintenance should I actually expect?","Measured across our own production groups, a storefront with no major redesign needs attention a handful of times a year, and a redesign costs a few hours per affected site. The trap is not the hours. It is that they arrive unannounced, all at once, and usually on the day something else is already on fire.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":183,"next_step":188},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.","12 min",false,{"title":75,"description":76,"keywords":77},"How to Debug a Broken Web Scraper: A Diagnosis Routine","A repeatable routine for diagnosing scraper failures: status before body, body before selector, one URL before the whole run. Includes the three false conclusions that waste the most time.","debug web scraper, scraper error, scraper returns nothing, scraper troubleshooting, http 403 scraper, scraper debugging",[79,84,107,112,118,148,166,177],{"type":80,"paragraphs":81},"prose",[82,83],"The instinct when a run comes back empty is to open the page in a browser, see the price sitting there in plain sight, and conclude that the selector is broken. That conclusion is right perhaps a third of the time, and the other two thirds are expensive, because you will rewrite a selector that was never wrong and the run will stay broken.","The routine below is ordered so that each step can only be reached if the previous one has been eliminated. It is deliberately mechanical. The whole point is to remove the guessing.",{"type":85,"title":86,"intro":87,"items":88},"steps","The routine","Do these in order on a single URL. Never debug against the whole run — you cannot see anything in five thousand rows of log.",[89,92,95,98,101,104],{"title":90,"text":91},"Pick one failing URL and work only on that","One URL that definitely used to work. If you cannot name one, that is itself the finding: you may be looking at content drift rather than breakage.",{"title":93,"text":94},"Check the status code before anything else","A 403, 429 or 503 ends the investigation. That is a bot wall or a rate limit, and nothing in your extraction layer is involved. Jump to lesson four.",{"title":96,"text":97},"Check what the body actually is","A 200 is not proof you got the product page. Challenge pages, consent interstitials, geo-redirects and soft 404s all return 200 with a perfectly valid HTML body. Look at the title tag and the length. A product page that is suddenly 4 KB is not a product page.",{"title":99,"text":100},"Search the raw body for the value, not for your selector","Take the price you can see in the browser and grep the fetched body for the digits. If the number is in there, you have an extraction problem. If it is not, the page you fetched is not the page you looked at, and that is a rendering or routing problem.",{"title":102,"text":103},"Only now look at the selector","You have earned the right to. And at this point the fix is usually obvious, because you know the value is present and you know where.",{"title":105,"text":106},"Re-run the single URL, then ten, then the group","Fixing one and immediately launching five thousand is how a wrong fix becomes an expensive wrong fix.",{"type":108,"variant":109,"title":110,"text":111},"callout","note","Grep the body before you theorise","Searching the fetched HTML for the literal digits of the price takes ten seconds and splits the problem space in half every time. It is the single highest-value habit in this course. We have skipped it, shipped a wrong fix on the strength of a plausible theory, and had to ship a second fix to undo the first.",{"type":80,"title":113,"paragraphs":114},"Three false conclusions this routine is built to prevent",[115,116,117],"The first is \"it works in my browser, so the site is fine\". Your browser has cookies, a residential IP, a full JavaScript engine and a history with that domain. It is not a control group. A page that loads for you and refuses a datacenter address is the ordinary case, not an anomaly — and it means a failure you cannot reproduce locally is still real.","The second is \"it returned 200, so we got the page\". Covered above, and it is the one that burns the most hours, because a 200 feels conclusive.","The third is subtler: \"the fix worked, the run is green\". Green after a fix means the run completed. It does not mean the rows are right. Compare the row count to last week's before you call it done — a selector that now matches a different element will happily return five thousand rows of the wrong thing.",{"type":119,"title":120,"headers":121,"rows":125},"table","Symptom to cause, once you have the evidence",[122,123,124],"Evidence","Cause","Where to go",[126,130,133,137,140,144],[127,128,129],"403 / 429 / 503","Bot wall or rate limit","Lesson 4",[131,132,129],"200, tiny body, unexpected title","Challenge or consent interstitial",[134,135,136],"200, full body, value absent from source","Rendered client-side","Lesson 3",[138,139,136],"200, value present in body, selector misses","Redesign",[141,142,143],"Some URLs fine, some 404","Content drift","Refresh the URL list",[145,146,147],"Everything green, fewer rows","Silent partial failure","Lesson 5",{"type":85,"title":149,"intro":150,"items":151},"Worked example: ten minutes on one URL","A feed that wrote 1,964 rows from 2,667 attempted pages. The temptation is to theorise — delisted products, a bad selector, a blocked range — and every theory sounds plausible. The routine answers it without a theory, and no step costs more than two minutes.",[152,155,158,161,163],{"title":153,"text":154},"Check the status distribution across the whole run, not a sample","If the failures all carry one status class you have the answer in ninety seconds. In this case every failure was a page-level error and not one was an empty result, which already kills the delisted-products theory: a delisted page returns something, it does not fail.",{"title":156,"text":157},"Fetch one failing URL and read the status","A 403 is a bot wall. A 404 is a stale URL list. A 200 means the cause is further down the page and the next two steps are worth the time.",{"title":159,"text":160},"Search the raw body for the literal value","Before looking at any selector, search the response for the price exactly as it appears on screen. Present means a selector problem. Absent means a rendering problem, and no amount of selector work will ever fix it.",{"title":102,"text":162},"Three of the four candidate causes have been eliminated for the price of four fetches. This is the point at which editing the selector is an informed act rather than a guess.",{"title":164,"text":165},"Prove the path ran before explaining why it failed","If a fix was deployed and nothing changed, establish that the new code executed at all — a timing difference, a log line, a billing delta at the provider. Explaining the failure of a code path that never ran is the most expensive way to spend an afternoon that is available to anybody.",{"type":167,"title":168,"intro":169,"items":170},"list","What usually goes wrong","Each of these turns a ten-minute diagnosis into a two-day one.",[171,172,173,174,175,176],"Theorising before fetching. Delisted products, a supplier feed change and a broken selector all explain the same row count, and a single fetch separates them.","Reading the status and stopping there. A 200 routinely carries a consent interstitial, a geo-redirect or a challenge page, all of which parse to nothing at all.","Testing from your own machine and concluding the site is fine. Your laptop has cookies, history and a residential address; the job has a flagged data centre range. The site treats them as different visitors, because they are.","Sampling the failures. The status distribution across the whole run is one query and it frequently ends the investigation by itself.","Changing three things and redeploying. When it works nobody knows which change worked, and when it does not nobody knows which one to undo.","Explaining a failure in code that never executed. Prove it ran — timing, logs, or a billing delta — before explaining anything about it.",{"type":80,"title":178,"paragraphs":179},"Prove the path ran before you explain why it failed",[180,181,182],"This is the rule that took us longest to learn and it generalises past scraping. Before building any theory about why a step failed, establish that the step executed at all.","Timing is the cheapest evidence: a fetch that supposedly hit a renderer and returned in 120 milliseconds did not hit a renderer. Billing is the next cheapest, if your provider exposes a credit balance — make one known-good call, note the delta, then diff the balance around the request you are investigating. If the counter did not move, the code path you are theorising about never ran, and every theory about its behaviour is noise.","We shipped one wrong fix and one wrong explanation in a single session by skipping that check. Both were plausible. Neither touched the actual cause.",[184,185,186,187],"Status, then body, then value-in-source, then selector. In that order, on one URL.","A 200 does not mean you got the page you asked for.","Your browser is not a control group — it has cookies, a residential IP and history.","Prove the code path executed before explaining why it misbehaved. Timing and billing deltas are the cheapest proof.",{"text":189,"label":190,"url":191},"Next: the extraction side — which selectors survive a rewrite and which are built to fail.","Lesson 3: selectors that survive a redesign","/learn/keep-scrapers-alive/selectors-that-survive-a-redesign",1,[194,200,201,207,212,217],{"slug":195,"navTitle":196,"title":197,"summary":198,"time":199,"needsAccount":73},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.","10 min",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":202,"navTitle":203,"title":204,"summary":205,"time":206,"needsAccount":73},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.","11 min",{"slug":208,"navTitle":209,"title":210,"summary":211,"time":206,"needsAccount":73},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":213,"navTitle":214,"title":215,"summary":216,"time":206,"needsAccount":73},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":218,"navTitle":219,"title":220,"summary":221,"time":199,"needsAccount":73},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"slug":195,"navTitle":196,"title":197,"summary":198,"time":199,"needsAccount":73},{"slug":202,"navTitle":203,"title":204,"summary":205,"time":206,"needsAccount":73},1791047867231]