[{"data":1,"prerenderedAt":239},["ShallowReactive",2],{"learn-lesson-keep-scrapers-alive-decide-what-to-do-when-a-site-wins":3},{"course":4,"lesson":67,"index":207,"outline":208,"prev":237,"next":238},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"keep-scrapers-alive",4,"You already have something running","6 lessons, about 65 minutes","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Keep Scrapers Alive: A Free 6-Lesson Maintenance Course","Why web scrapers break and what to do about it. Reading failures properly, selectors that survive redesigns, bot walls, monitoring the feed rather than the run, and knowing when to stop. Free, ungated.","web scraper maintenance, scraper broke, scraper monitoring, css selector best practice, bot detection, scraper error handling, data quality monitoring, web scraping reliability","A free course on why scrapers break and how to keep them running","Six written lessons on scraper maintenance: failure diagnosis, durable selectors, bot walls, feed monitoring, and deciding when a site has won.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course four","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.",{"label":21,"url":22},"Start with lesson one","/learn/keep-scrapers-alive/why-scrapers-break",{"label":24,"url":25},"See real run output","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Name the five reasons a scraper stops working, and tell them apart from the evidence rather than from a hunch","Read a failure down to its cause instead of restarting the job and hoping","Write selectors that survive a front-end rewrite, and recognise the ones that will not","Understand what a bot wall is actually measuring, and what genuinely changes the outcome","Monitor the output of a feed, not just whether the job exited zero","Decide, with a rule rather than a mood, when to stop fighting a site",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Anyone who already has scrapers running and keeps getting surprised by them","Developers who inherited somebody else's collection layer and want to stop firefighting it","Analysts who depend on a scraped feed and need to know how much to trust it","Not written for",[44,45],"First-time scraper authors — start with course one or course three, this assumes something already runs","Anyone looking for techniques to defeat a specific site's protection. That is not what this is, and lesson four explains why that framing loses.",{"title":47,"intro":48},"The six lessons","Read in order the first time. After that it works as a reference — lesson two is the one people come back to.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","What people usually want to know when a scraper has just broken.",[54,57,60,63],{"title":55,"description":56},"My scraper broke this morning. Which lesson do I read?","Lesson two. It is the diagnosis lesson and it is deliberately the longest. Most wasted maintenance time comes from fixing the wrong thing — rewriting a selector when the page never loaded, or rotating proxies when the product was simply discontinued.",{"title":58,"description":59},"Is this about avoiding detection?","No. Lesson four covers bot walls because you will meet them, but the framing there is about what the wall is measuring and what honestly changes the outcome, not about evasion. The most reliable answers turn out to be boring: slow down, be identifiable, and accept that some sites will say no.",{"title":61,"description":62},"Does any of this apply if I use a managed service?","The diagnosis and monitoring lessons especially. A vendor can own the fetching and the extraction, but nobody else can know that your category page should have returned two thousand rows and came back with four hundred. That judgement stays with you whatever you buy.",{"title":64,"description":65},"How much maintenance should I actually expect?","Measured across our own production groups, a storefront with no major redesign needs attention a handful of times a year, and a redesign costs a few hours per affected site. The trap is not the hours. It is that they arrive unannounced, all at once, and usually on the day something else is already on fire.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":198,"next_step":203},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.","10 min",false,{"title":75,"description":76,"keywords":77},"When to Stop Trying to Scrape a Website","A decision rule for repairing, routing around, or abandoning a source. How to value the data, cap the effort, and write a coverage gap report that is useful.","when to stop scraping, scraping coverage gap, data sourcing decision, alternative data source, scraper cost benefit",[79,84,104,110,130,135,142,178,189],{"type":80,"paragraphs":81},"prose",[82,83],"At some point a source becomes more expensive than it is worth, and the discipline is to notice that on purpose rather than by exhaustion. Teams that handle this badly do not usually make the wrong call — they make no call at all, and spend six months half-maintaining something nobody has decided to keep.","There are three available answers and it is worth being explicit about which one you are choosing.",{"type":85,"title":86,"headers":87,"rows":91},"table","The three answers",[88,89,90],"Choice","When it is right","What it costs",[92,96,100],[93,94,95],"Repair","Cause is identified, fix is bounded, source is load-bearing","Hours now, and the same hours again at the next redesign",[97,98,99],"Route around","The same data exists somewhere less defended","A different source means different coverage — say so",[101,102,103],"Stop","Cost per page exceeds the value, or the block is categorical","A documented gap, which is a real cost, honestly priced",{"type":80,"title":105,"paragraphs":106},"Put a number on the data before you argue about the effort",[107,108,109],"Most of these debates go badly because one side is talking about difficulty and the other about importance, and neither has been quantified.","The question that resolves it: if this source disappeared tomorrow, what decision gets made worse? If the answer is \"our weekly price position on four hundred SKUs we actively compete on\", that is load-bearing and you should spend real effort. If the answer is \"a column on a dashboard that two people open\", you have your decision and it is not the one anyone was arguing for.","Then cap the effort before starting. A day, a week, whatever is proportionate — decided in advance, because the sunk cost will absolutely argue for one more afternoon and it will say that every afternoon.",{"type":111,"title":112,"intro":113,"items":114},"steps","The routing-around checklist","Before concluding a source is unavailable, the same data is often published somewhere nobody thought to look.",[115,118,121,124,127],{"title":116,"text":117},"The site's own feed","Sitemaps, product feeds, affiliate exports and RSS are published deliberately and defended far less than the HTML.",{"title":119,"text":120},"A marketplace listing","Plenty of retailers that block you directly also sell on a marketplace that does not, with the same prices attached.",{"title":122,"text":123},"A comparison site or aggregator","Lower precision and a lag, but a legitimate second-best, provided you label the provenance.",{"title":125,"text":126},"The mobile application's backend","Often a clean JSON API. Check the terms before relying on it — this is where the legal question stops being theoretical, and course six covers it.",{"title":128,"text":129},"Asking","Genuinely underused. Some retailers will hand you a feed if you explain what you need and why. A no costs you an email.",{"type":131,"variant":132,"title":133,"text":134},"callout","note","Label the substitute, always","If prices for one retailer come from a marketplace listing rather than their own site, that belongs in the data as a provenance field, not in a comment in the code. Six months later nobody will remember, and a slightly different number will be read as a price change rather than a sourcing difference.",{"type":80,"title":136,"paragraphs":137},"How to write a coverage gap so it is useful",[138,139,140,141],"A good gap report is three sentences and a number. What is missing, what was tried, what it would take, and what it costs to leave it.","\"We cannot read this retailer. Nine hundred pages attempted over two days, all refused at the network level before any content was returned; the pattern is categorical rather than rate-related. Getting through would mean a per-page cost roughly four times our current average, with no guarantee of durability. Leaving it means our price position excludes one of eleven competitors in that segment, which matters most on garden furniture where they are the volume leader.\"","That is an input to a decision. Compare it to \"the scraper for this site doesn't work\", which is an apology and tells nobody anything.","The habit underneath all of it: publish what you measured, not what you expected. A page that honestly says the storefront did not return readable product data is worth more than one that pads the gap with a plausible number, because the first can be acted on and the second quietly poisons everything downstream of it.",{"type":85,"title":143,"intro":144,"headers":145,"rows":147},"Worked example: pricing the three answers on one site","A competitor that made up 11% of the feed has gone behind a protection layer. Laid out like this the decision takes an hour. Left undocumented it takes a quarter, and the answer arrived at is usually the one that is not in this table.",[146,93,97,101],"",[148,153,158,163,168,173],[149,150,151,152],"What it means here","Escalate to rendering plus a residential address","Take the price from the marketplace listing instead","Remove the site and document the gap",[154,155,156,157],"Effort","about two days, then ongoing","half a day","an hour",[159,160,161,162],"Running cost","about 25× the base page rate, on 11% of the feed","base rate, different source","nothing",[164,165,166,167],"What the data becomes","unchanged","a third-party seller's price, not the retailer's — label it","a named absence",[169,170,171,172],"The risk you are taking","it breaks again at the next change, now at 25×","the substitute gets treated as equivalent by everyone downstream","somebody presents a market view with a hole in it",[174,175,176,177],"When it wins","the site sets the market price on your top lines","the substitute is genuinely comparable","the site was a benchmark rather than a threat",{"type":179,"title":180,"intro":181,"items":182},"list","What usually goes wrong","The three answers are all defensible. What costs money is the fourth thing people do instead.",[183,184,185,186,187,188],"Choosing none of them. The default is a scraper half-fixed every few weeks forever, which costs more than any single column in the table.","Arguing about effort before valuing the data. The question is which decision gets worse without this site, and if the answer is none, the right-hand column is free.","Routing around without labelling the substitute. A marketplace price and a retailer price are different things, and once they share a column nobody can separate them again.","Escalating permanently to solve a temporary problem. The expensive rung tends to stay switched on long after the reason for switching it on has gone.","Writing the coverage gap as an apology. \"It doesn't work\" is not a decision input. \"This competitor is 11% of the feed, unavailable since 14 March, substitute available at marketplace level\" is.","Scheduling no review. A gap documented once and never revisited becomes permanent through neglect rather than through a decision.",{"type":179,"title":190,"intro":191,"items":192},"The maintenance routine worth having","None of this is clever. All of it compounds.",[193,194,195,196,197],"Review the row-count deltas weekly. Ten minutes, and it is the whole early warning system.","Refresh URL lists monthly — content drift is constant and silent.","Hand-check twenty rows against live pages monthly.","Keep a one-line log per source of what broke and what fixed it. It is how a new person becomes useful in a week instead of a quarter.","Re-run the value question annually. Sources that were load-bearing stop being load-bearing and nobody notices until someone audits the bill.",[199,200,201,202],"Three answers: repair, route around, stop. Choosing none of them is the expensive default.","Value the data first — what decision gets worse without it — then cap the effort before you start.","Sitemaps, marketplaces, aggregators, mobile backends and simply asking are all real alternatives. Label the substitute in the data.","A specific coverage gap with numbers is a decision input. \"It doesn't work\" is an apology.",{"text":204,"label":205,"url":206},"Collected data is only useful once rows from different sites can be matched to the same product. That is the next course.","Course 5: matching products across sites","/learn/matching-products-across-sites",5,[209,214,220,226,231,236],{"slug":210,"navTitle":211,"title":212,"summary":213,"time":72,"needsAccount":73},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.",{"slug":215,"navTitle":216,"title":217,"summary":218,"time":219,"needsAccount":73},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.","12 min",{"slug":221,"navTitle":222,"title":223,"summary":224,"time":225,"needsAccount":73},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.","11 min",{"slug":227,"navTitle":228,"title":229,"summary":230,"time":225,"needsAccount":73},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":232,"navTitle":233,"title":234,"summary":235,"time":225,"needsAccount":73},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":232,"navTitle":233,"title":234,"summary":235,"time":225,"needsAccount":73},null,1791047867308]