[{"data":1,"prerenderedAt":231},["ShallowReactive",2],{"learn-lesson-keep-scrapers-alive-bot-walls-and-what-actually-works":3},{"course":4,"lesson":67,"index":199,"outline":200,"prev":229,"next":230},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"keep-scrapers-alive",4,"You already have something running","6 lessons, about 65 minutes","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Keep Scrapers Alive: A Free 6-Lesson Maintenance Course","Why web scrapers break and what to do about it. Reading failures properly, selectors that survive redesigns, bot walls, monitoring the feed rather than the run, and knowing when to stop. Free, ungated.","web scraper maintenance, scraper broke, scraper monitoring, css selector best practice, bot detection, scraper error handling, data quality monitoring, web scraping reliability","A free course on why scrapers break and how to keep them running","Six written lessons on scraper maintenance: failure diagnosis, durable selectors, bot walls, feed monitoring, and deciding when a site has won.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course four","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.",{"label":21,"url":22},"Start with lesson one","/learn/keep-scrapers-alive/why-scrapers-break",{"label":24,"url":25},"See real run output","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Name the five reasons a scraper stops working, and tell them apart from the evidence rather than from a hunch","Read a failure down to its cause instead of restarting the job and hoping","Write selectors that survive a front-end rewrite, and recognise the ones that will not","Understand what a bot wall is actually measuring, and what genuinely changes the outcome","Monitor the output of a feed, not just whether the job exited zero","Decide, with a rule rather than a mood, when to stop fighting a site",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Anyone who already has scrapers running and keeps getting surprised by them","Developers who inherited somebody else's collection layer and want to stop firefighting it","Analysts who depend on a scraped feed and need to know how much to trust it","Not written for",[44,45],"First-time scraper authors — start with course one or course three, this assumes something already runs","Anyone looking for techniques to defeat a specific site's protection. That is not what this is, and lesson four explains why that framing loses.",{"title":47,"intro":48},"The six lessons","Read in order the first time. After that it works as a reference — lesson two is the one people come back to.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","What people usually want to know when a scraper has just broken.",[54,57,60,63],{"title":55,"description":56},"My scraper broke this morning. Which lesson do I read?","Lesson two. It is the diagnosis lesson and it is deliberately the longest. Most wasted maintenance time comes from fixing the wrong thing — rewriting a selector when the page never loaded, or rotating proxies when the product was simply discontinued.",{"title":58,"description":59},"Is this about avoiding detection?","No. Lesson four covers bot walls because you will meet them, but the framing there is about what the wall is measuring and what honestly changes the outcome, not about evasion. The most reliable answers turn out to be boring: slow down, be identifiable, and accept that some sites will say no.",{"title":61,"description":62},"Does any of this apply if I use a managed service?","The diagnosis and monitoring lessons especially. A vendor can own the fetching and the extraction, but nobody else can know that your category page should have returned two thousand rows and came back with four hundred. That judgement stays with you whatever you buy.",{"title":64,"description":65},"How much maintenance should I actually expect?","Measured across our own production groups, a storefront with no major redesign needs attention a handful of times a year, and a redesign costs a few hours per affected site. The trap is not the hours. It is that they arrive unannounced, all at once, and usually on the day something else is already on fire.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":190,"next_step":195},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.","11 min",false,{"title":75,"description":76,"keywords":77},"Why Scrapers Get Blocked: Bot Detection Explained","What bot detection actually measures, why a flagged datacenter IP is refused rather than challenged, and which responses genuinely change the outcome. Plus when to stop.","bot detection, scraper blocked 403, cloudflare scraper, rate limiting scraper, datacenter ip blocked, residential proxy, web scraping blocked",[79,84,112,117,125,130,140,175,185],{"type":80,"paragraphs":81},"prose",[82,83],"A protection layer is not trying to work out whether you are a robot. It is scoring how much you cost and how much you are worth, and the score is mostly made of three things: where you are connecting from, how fast you are asking, and whether your request looks like it came from the software it claims to come from.","That framing matters, because it tells you which of your options are real. You cannot argue with the score. You can change the inputs to it, and only some of those inputs are yours to change.",{"type":85,"title":86,"headers":87,"rows":91},"table","What gets measured, and what you can do about it",[88,89,90],"Signal","What it sees","Your realistic lever",[92,96,100,104,108],[93,94,95],"IP reputation","Datacenter range, known provider, prior abuse from that block","Route differently, or accept the refusal",[97,98,99],"Request rate","Requests per minute from one source, and the regularity of the gaps","Slow down and vary. This is the biggest free win.",[101,102,103],"Header coherence","Whether your headers describe a real client consistently","Send a coherent, honest set. Do not impersonate.",[105,106,107],"Session behaviour","Straight to product pages, no cookies, no referrer, never a listing","Behave like a reader: accept cookies, follow links",[109,110,111],"Browser fingerprint","JavaScript challenges, canvas, timing","Render properly, or stop",{"type":113,"variant":114,"title":115,"text":116},"callout","warning","The laptop test will lie to you","A flagged datacenter address is often refused outright, not challenged. So the body fingerprint you carefully measured from your own machine — the challenge page, the specific wording, the script tag you keyed your detection on — never appears on the deployment host, which simply gets a 403. Any detection logic built from a local reproduction is a false negative in production. Key your gate on the HTTP status, which is the one signal that survives the move.",{"type":80,"title":118,"paragraphs":119},"The boring answers outperform the clever ones",[120,121,122,123,124],"Nearly everything written about this topic is about defeating detection, and nearly all of it ages badly, because it describes a specific countermeasure against a specific version of a specific vendor's product. The things that keep working are unglamorous.","Ask less often. Most scraping is dramatically over-frequent relative to how often the underlying data changes. A price that moves twice a week does not need hourly polling, and halving your request rate removes the signal you were most visibly failing on.","Be identifiable. A user agent that says who you are and a contact address is not naivety — it is the thing that gets you an allowlist instead of a ban when somebody looks at the logs. Pretending to be Chrome and failing the fingerprint check is worse than both.","Ask for less. Pulling a thousand product pages when the category listing already carries the price and the title is a thousand requests you did not need, and a hundredfold increase in your visibility.","Cache aggressively. The cheapest request is the one you did not make.",{"type":80,"title":126,"paragraphs":127},"One measured thing about market and locale",[128,129],"Related, and it has cost us real money. Large international retailers decide which market you are in, and that decision changes the price you are served — not the formatting, the actual number.","On Amazon we measured this directly. Sending an Accept-Language header alone did nothing at all: three attempts within the same minute, same market, same price every time. The cookie is what flips the market. Get that wrong and your scraper does not fail — it succeeds, and returns a correct price from the wrong country, which is the silent-partial failure from lesson one wearing a different coat. A single-currency sanity check on the output would have caught it in a day. We did not have one, and the wrong number was live for longer than it should have been.",{"type":131,"title":132,"intro":133,"items":134},"list","When to stop","Some sites will not be read, and continuing is a cost with no return. Stop when any two of these are true.",[135,136,137,138,139],"You have spent more than a day on one site and the success rate is still under half","The fix involves solving a visual or interactive challenge","You are being asked to log in, and the terms attached to that login forbid automated access","The per-page cost of getting through exceeds what the data is worth to you","The thing you need is published somewhere else — a feed, a marketplace listing, an aggregator — that is not fighting you",{"type":85,"title":141,"intro":142,"headers":143,"rows":148},"Worked example: the escalation ladder, priced","Each rung costs more than the one above it, and the order matters because the cheap rungs resolve most cases. Cost is given as a multiple of a plain fetch rather than in currency, because the absolute figures differ by provider and by month while the ratios are the part that generalises — and the ratios are what should govern the decision.",[144,145,146,147],"Rung","What it changes","Relative cost","When it is the right answer",[149,153,156,161,165,170],[150,97,151,152],"Slow down","×1","Almost always first — rate is the cheapest thing you control",[154,101,151,155],"Fix the headers","A client whose headers match no real browser",[157,158,159,160],"Render the page","Execution of page scripts","about ×5","The value is absent from the source and present on screen",[162,93,163,164],"Residential address","about ×10","A data centre range is being refused outright",[166,167,168,169],"Both together","Both","about ×25","Rarely, and never without a spend cap in front of it",[171,172,173,174],"Stop","Nothing","×0","A documented coverage gap — the next lesson is about writing it",{"type":131,"title":176,"intro":177,"items":178},"What usually goes wrong","Almost every expensive mistake here is an escalation made before the cheap rungs were tried, or measured somewhere the job does not run.",[179,180,181,182,183,184],"Starting at the expensive rung. Residential addresses and full rendering fix a minority of cases and cost five to twenty-five times as much on every page, permanently.","Reproducing it on a laptop. A flagged data centre address is refused outright rather than challenged, so a body fingerprint measured from your machine is a false negative on the host that actually runs the job. Key the detection on the HTTP status.","Escalating without a spend cap. The combined rung is roughly twenty-five times the base rate, and a retry loop on top of it is how a month's budget disappears in an afternoon.","Changing the market without meaning to. On at least one large retailer the cookie decides the market and the language header alone does nothing, so an escalation can return a perfectly correct price from the wrong country.","Impersonating a browser as a strategy. It is an arms race against a team whose whole job it is, and it converts a conduct question into a bad-faith one — course six covers what that costs.","Treating stopping as a failure. A documented coverage gap with numbers attached is a result; a partial feed nobody labelled is a liability.",{"type":80,"title":186,"paragraphs":187},"Stopping is a result",[188,189],"A coverage note saying \"this retailer is not readable, here is the evidence, here is what we propose instead\" is a legitimate deliverable. It is worth considerably more than a feed that is quietly thirty percent wrong because somebody refused to be beaten.","The failure mode to avoid is the one where a site that mostly blocks you produces a trickle of rows and nobody labels them. A partial feed presented as a complete one is the most expensive artefact in this whole discipline.",[191,192,193,194],"Detection scores IP reputation, request rate, header coherence and session behaviour. Rate is the one you control most cheaply.","A flagged datacenter IP is refused, not challenged — detection logic built from a laptop reproduction is a false negative in production. Key on HTTP status.","Measured on Amazon: the cookie flips the market, Accept-Language alone does nothing. Wrong market means a correct price from the wrong country.","Stopping with a documented coverage gap beats a partial feed nobody labelled.",{"text":196,"label":197,"url":198},"Next: the failure that does not announce itself — monitoring output rather than exit codes.","Lesson 5: monitor the feed, not the run","/learn/keep-scrapers-alive/monitor-the-feed-not-the-run",3,[201,207,213,218,219,224],{"slug":202,"navTitle":203,"title":204,"summary":205,"time":206,"needsAccount":73},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.","10 min",{"slug":208,"navTitle":209,"title":210,"summary":211,"time":212,"needsAccount":73},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.","12 min",{"slug":214,"navTitle":215,"title":216,"summary":217,"time":72,"needsAccount":73},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":220,"navTitle":221,"title":222,"summary":223,"time":72,"needsAccount":73},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":225,"navTitle":226,"title":227,"summary":228,"time":206,"needsAccount":73},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"slug":214,"navTitle":215,"title":216,"summary":217,"time":72,"needsAccount":73},{"slug":220,"navTitle":221,"title":222,"summary":223,"time":72,"needsAccount":73},1791047867259]