[{"data":1,"prerenderedAt":223},["ShallowReactive",2],{"learn-lesson-keep-scrapers-alive-why-scrapers-break":3},{"course":4,"lesson":67,"index":191,"outline":192,"prev":221,"next":222},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"keep-scrapers-alive",4,"You already have something running","6 lessons, about 65 minutes","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Keep Scrapers Alive: A Free 6-Lesson Maintenance Course","Why web scrapers break and what to do about it. Reading failures properly, selectors that survive redesigns, bot walls, monitoring the feed rather than the run, and knowing when to stop. Free, ungated.","web scraper maintenance, scraper broke, scraper monitoring, css selector best practice, bot detection, scraper error handling, data quality monitoring, web scraping reliability","A free course on why scrapers break and how to keep them running","Six written lessons on scraper maintenance: failure diagnosis, durable selectors, bot walls, feed monitoring, and deciding when a site has won.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course four","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.",{"label":21,"url":22},"Start with lesson one","/learn/keep-scrapers-alive/why-scrapers-break",{"label":24,"url":25},"See real run output","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Name the five reasons a scraper stops working, and tell them apart from the evidence rather than from a hunch","Read a failure down to its cause instead of restarting the job and hoping","Write selectors that survive a front-end rewrite, and recognise the ones that will not","Understand what a bot wall is actually measuring, and what genuinely changes the outcome","Monitor the output of a feed, not just whether the job exited zero","Decide, with a rule rather than a mood, when to stop fighting a site",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Anyone who already has scrapers running and keeps getting surprised by them","Developers who inherited somebody else's collection layer and want to stop firefighting it","Analysts who depend on a scraped feed and need to know how much to trust it","Not written for",[44,45],"First-time scraper authors — start with course one or course three, this assumes something already runs","Anyone looking for techniques to defeat a specific site's protection. That is not what this is, and lesson four explains why that framing loses.",{"title":47,"intro":48},"The six lessons","Read in order the first time. After that it works as a reference — lesson two is the one people come back to.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","What people usually want to know when a scraper has just broken.",[54,57,60,63],{"title":55,"description":56},"My scraper broke this morning. Which lesson do I read?","Lesson two. It is the diagnosis lesson and it is deliberately the longest. Most wasted maintenance time comes from fixing the wrong thing — rewriting a selector when the page never loaded, or rotating proxies when the product was simply discontinued.",{"title":58,"description":59},"Is this about avoiding detection?","No. Lesson four covers bot walls because you will meet them, but the framing there is about what the wall is measuring and what honestly changes the outcome, not about evasion. The most reliable answers turn out to be boring: slow down, be identifiable, and accept that some sites will say no.",{"title":61,"description":62},"Does any of this apply if I use a managed service?","The diagnosis and monitoring lessons especially. A vendor can own the fetching and the extraction, but nobody else can know that your category page should have returned two thousand rows and came back with four hundred. That judgement stays with you whatever you buy.",{"title":64,"description":65},"How much maintenance should I actually expect?","Measured across our own production groups, a storefront with no major redesign needs attention a handful of times a year, and a redesign costs a few hours per affected site. The trap is not the hours. It is that they arrive unannounced, all at once, and usually on the day something else is already on fire.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":182,"next_step":187},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.","10 min",false,{"title":75,"description":76,"keywords":77},"Why Web Scrapers Break: The Five Real Causes","Redesigns, bot walls, rendering changes, content drift and silent partial failure. What each one looks like, how often it happens, and why the fixes are different.","why scrapers break, web scraper maintenance, scraper stopped working, scraper reliability, data pipeline failure",[79,84,119,125,130,140,168,177],{"type":80,"paragraphs":81},"prose",[82,83],"A scraper that worked yesterday and does not work today has failed for one of five reasons. They look similar from the outside — an empty result, or fewer rows than usual — and they have almost nothing in common underneath. The habit worth building is naming the category before touching any code, because four of the five fixes are wrong four fifths of the time.","In rough order of how often they occur on a mature setup:",{"type":85,"title":86,"intro":87,"headers":88,"rows":93},"table","The five failure modes","Frequency here is from our own production groups, not an industry figure. Your mix will differ by sector; the categories will not.",[89,90,91,92],"Failure","What you see","What actually changed","Typical fix",[94,99,104,109,114],[95,96,97,98],"Content drift","Fewer rows, no errors","Products delisted, renamed or moved; the site is fine","Nothing. Update the URL list.",[100,101,102,103],"Redesign","Zero rows, page fetched fine","The markup moved; your selector points at nothing","Re-point the selector",[105,106,107,108],"Bot wall","Non-200 status, or a page that is not the product page","The site decided you are not a browser","Slow down, identify yourself, or route differently",[110,111,112,113],"Rendering change","Page fetched, markup present, values empty","Content moved behind JavaScript","Render the page, or read the data source directly",[115,116,117,118],"Silent partial","Looks normal, is wrong","Pagination capped, a region defaulted, a currency flipped","The hardest. Lesson five.",{"type":80,"title":120,"paragraphs":121},"The one that costs the most is not the one you expect",[122,123,124],"Redesigns feel like the enemy because they are dramatic — everything goes to zero and somebody notices within the hour. That visibility is a gift. A loud failure is a cheap failure.","The expensive category is the last row. A run that returns four hundred rows where it used to return two thousand exits successfully. The scheduler is happy. The dashboard has data in it. And every number downstream is now computed over a fifth of the catalogue, which is worse than having no number at all, because somebody will price against it.","We have measured exactly this on a live client group: rows per week fell from 10,382 to 8,226 across two weeks while the total number of requests stayed identical — 11,032 then 11,034. Nothing errored. Nothing alerted. The job had been fetching and billing for every page the whole time and simply returning less.",{"type":126,"variant":127,"title":128,"text":129},"callout","warning","Exit code zero is not a health check","The single most common monitoring mistake is treating \"the job finished\" as \"the job worked\". Those are different claims and only one of them is worth an alert. Lesson five is entirely about the difference, and it is the lesson that changes the most for most teams.",{"type":131,"title":132,"intro":133,"items":134},"list","Questions that place a failure in the right category in under a minute","Ask these in order. The first one that gives an interesting answer is usually the whole diagnosis.",[135,136,137,138,139],"What HTTP status came back? A non-200 is a bot wall or an outage, and no selector work will help.","Did the page body arrive, and is it the page you asked for? A 200 that returns a challenge page is still a block.","Is the field present in the markup but empty, or absent entirely? Present-but-empty points at rendering; absent points at a redesign.","Did every URL fail, or a subset? A subset is almost always content drift.","Did the row count fall, or go to zero? Zero is loud and simple. A fall is the dangerous one.",{"type":85,"title":141,"intro":142,"headers":143,"rows":147},"Worked example: one week of symptoms, placed in five minutes","Five scrapers, five symptoms, five different fixes. The column to read is the middle one: the evidence that separates the categories is cheap to collect, and almost nobody collects it, which is why maintenance feels like one endless problem rather than five bounded ones.",[144,145,146],"Symptom","The evidence that settles it","Category, and what the fix is",[148,152,156,160,164],[149,150,151],"Zero rows, job red, connection refused","HTTP status on a single URL","Bot wall — slow down, or stop and document the gap",[153,154,155],"Zero rows, job green, 200 returned","Is the price in the HTML source, or only on screen?","Rendering change — the page went client-side",[157,158,159],"Price null on every row, everything else intact","Search the response body for the literal price string","Content drift — one selector, five minutes",[161,162,163],"Every field null, 200 returned, body unrecognisable","Open the page in a browser","Redesign — rewrite the extraction for this site",[165,166,167],"Rows down 21% over two weeks, no errors anywhere","Row count against the trailing average, per site","Silent partial — the expensive one",{"type":131,"title":169,"intro":170,"items":171},"What usually goes wrong","Four of these five are attempts to fix a scraper before anyone has said which of the five things is wrong with it.",[172,173,174,175,176],"Treating all five as one problem called maintenance. They have four different fixes and one of them is \"stop\", so the category has to be named before anyone touches code.","Reaching for proxies first. It is the standard answer to one of the five modes, the wrong answer to the other four, and expensive in all five.","Using exit code zero as the health signal. The worst mode in the table above returns successfully, every day, for a fortnight.","Diagnosing from a browser. Yours carries cookies, history and a residential address; the job has none of them, which is why the page you are looking at is not the page it received.","Rewriting the selector before checking whether the value is in the source at all. If the price only exists after JavaScript runs, no selector will ever find it, and you can spend a day proving that.",{"type":80,"title":178,"paragraphs":179},"Why it matters that these are five and not one",[180,181],"Because the reflex fix for each is useless for the other four. Rotating proxies does nothing about a redesign. Rewriting a selector does nothing about a product that no longer exists. Adding a browser renderer to a site that was never JavaScript-driven costs you money per page and fixes nothing.","Most teams that describe scraping as \"constant maintenance\" are not doing more maintenance than anyone else. They are doing the same amount, spent on the wrong category, which means the actual cause survives the fix and comes back next week.",[183,184,185,186],"Five failure modes: content drift, redesign, bot wall, rendering change, silent partial.","Loud failures are cheap. The expensive one returns successfully with less data.","A measured case: rows fell 21% across two weeks with request volume unchanged and zero errors.","Name the category before you touch code — four of the five standard fixes are wrong most of the time.",{"text":188,"label":189,"url":190},"Next: how to actually read a failure instead of guessing which of the five it was.","Lesson 2: read the failure, not the symptom","/learn/keep-scrapers-alive/read-the-failure-not-the-symptom",0,[193,194,200,206,211,216],{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":195,"navTitle":196,"title":197,"summary":198,"time":199,"needsAccount":73},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.","12 min",{"slug":201,"navTitle":202,"title":203,"summary":204,"time":205,"needsAccount":73},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.","11 min",{"slug":207,"navTitle":208,"title":209,"summary":210,"time":205,"needsAccount":73},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":212,"navTitle":213,"title":214,"summary":215,"time":205,"needsAccount":73},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":217,"navTitle":218,"title":219,"summary":220,"time":72,"needsAccount":73},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",null,{"slug":195,"navTitle":196,"title":197,"summary":198,"time":199,"needsAccount":73},1791047867195]