[{"data":1,"prerenderedAt":235},["ShallowReactive",2],{"learn-lesson-keep-scrapers-alive-monitor-the-feed-not-the-run":3},{"course":4,"lesson":67,"index":6,"outline":204,"prev":233,"next":234},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"keep-scrapers-alive",4,"You already have something running","6 lessons, about 65 minutes","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Keep Scrapers Alive: A Free 6-Lesson Maintenance Course","Why web scrapers break and what to do about it. Reading failures properly, selectors that survive redesigns, bot walls, monitoring the feed rather than the run, and knowing when to stop. Free, ungated.","web scraper maintenance, scraper broke, scraper monitoring, css selector best practice, bot detection, scraper error handling, data quality monitoring, web scraping reliability","A free course on why scrapers break and how to keep them running","Six written lessons on scraper maintenance: failure diagnosis, durable selectors, bot walls, feed monitoring, and deciding when a site has won.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course four","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.",{"label":21,"url":22},"Start with lesson one","/learn/keep-scrapers-alive/why-scrapers-break",{"label":24,"url":25},"See real run output","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Name the five reasons a scraper stops working, and tell them apart from the evidence rather than from a hunch","Read a failure down to its cause instead of restarting the job and hoping","Write selectors that survive a front-end rewrite, and recognise the ones that will not","Understand what a bot wall is actually measuring, and what genuinely changes the outcome","Monitor the output of a feed, not just whether the job exited zero","Decide, with a rule rather than a mood, when to stop fighting a site",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Anyone who already has scrapers running and keeps getting surprised by them","Developers who inherited somebody else's collection layer and want to stop firefighting it","Analysts who depend on a scraped feed and need to know how much to trust it","Not written for",[44,45],"First-time scraper authors — start with course one or course three, this assumes something already runs","Anyone looking for techniques to defeat a specific site's protection. That is not what this is, and lesson four explains why that framing loses.",{"title":47,"intro":48},"The six lessons","Read in order the first time. After that it works as a reference — lesson two is the one people come back to.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","What people usually want to know when a scraper has just broken.",[54,57,60,63],{"title":55,"description":56},"My scraper broke this morning. Which lesson do I read?","Lesson two. It is the diagnosis lesson and it is deliberately the longest. Most wasted maintenance time comes from fixing the wrong thing — rewriting a selector when the page never loaded, or rotating proxies when the product was simply discontinued.",{"title":58,"description":59},"Is this about avoiding detection?","No. Lesson four covers bot walls because you will meet them, but the framing there is about what the wall is measuring and what honestly changes the outcome, not about evasion. The most reliable answers turn out to be boring: slow down, be identifiable, and accept that some sites will say no.",{"title":61,"description":62},"Does any of this apply if I use a managed service?","The diagnosis and monitoring lessons especially. A vendor can own the fetching and the extraction, but nobody else can know that your category page should have returned two thousand rows and came back with four hundred. That judgement stays with you whatever you buy.",{"title":64,"description":65},"How much maintenance should I actually expect?","Measured across our own production groups, a storefront with no major redesign needs attention a handful of times a year, and a redesign costs a few hours per affected site. The trap is not the hours. It is that they arrive unannounced, all at once, and usually on the day something else is already on fire.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":195,"next_step":200},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.","11 min",false,{"title":75,"description":76,"keywords":77},"Data Quality Monitoring for Web Scrapers: Six Checks","How to detect a scraper that completes successfully but returns wrong data. Row count deltas, null rates, distribution shifts, staleness, and thresholds that avoid alert fatigue.","data quality monitoring, scraper monitoring, data pipeline alerting, silent data failure, row count anomaly detection, price data validation",[79,84,117,122,127,133,144,151,179,189],{"type":80,"paragraphs":81},"prose",[82,83],"Everything in this lesson exists because of one fact: the worst scraper failures do not raise errors. They complete, they write rows, the dashboard refreshes, and the numbers are wrong.","Monitoring the run tells you the job finished. Monitoring the feed tells you whether the thing it produced is usable. Only the second one is worth waking up for, and most teams have only built the first.",{"type":85,"title":86,"intro":87,"headers":88,"rows":92},"table","Six checks, cheapest first","Implement them in this order. The first two cover most of the real-world damage.",[89,90,91],"Check","Catches","Suggested trigger",[93,97,101,105,109,113],[94,95,96],"Row count versus trailing average","Pagination caps, partial blocks, truncated lists","Deviation beyond about 20% of the 7-day mean",[98,99,100],"Null rate per field","A single selector breaking while the rest hold","Any field whose null rate doubles week on week",[102,103,104],"Value distribution","Currency flips, unit errors, wrong market","Median moves more than 15% with no known cause",[106,107,108],"Staleness","A cached response being re-served as fresh","Any row whose value is byte-identical for an implausible stretch",[110,111,112],"Coverage against the expected URL list","Silent drops you never asked about","Attempted versus returned falls below your agreed floor",[114,115,116],"Cross-source agreement","Everything else, where you have a second source","Two sources disagreeing by more than a tolerance",{"type":80,"title":118,"paragraphs":119},"Row count is the single highest-value check",[120,121],"If you only ever build one of these, build this one. Store the count per source per run and compare it to the trailing average. It is a handful of lines and it catches the majority of silent failures, because almost every silent failure shows up first as \"less than usual\".","The live case from lesson one is exactly this shape. Rows per week went 10,382 then 8,237 then 8,226, while total requests stayed at 11,032, 11,035 and 11,034. Every page was fetched. Every page was billed. A fifth of the output simply stopped arriving, and because the error count was zero and the job was green, nothing anywhere said so. A row-count delta check would have flagged it in week one.",{"type":123,"variant":124,"title":125,"text":126},"callout","warning","Count what was attempted, not just what succeeded","A success rate computed over the rows you got is meaningless — it is always close to 100%, by construction. The denominator has to be what you intended to fetch. Attempted versus returned is the number that tells you the truth, and it is the number that is hardest to get after the fact, which is why it has to be recorded at run time.",{"type":80,"title":128,"paragraphs":129},"Thresholds, and the alert nobody reads",[130,131,132],"A monitor that fires every day is not a monitor, it is a background noise generator, and the second week of it is more dangerous than having no monitor at all because now everyone has learned to dismiss that channel.","Three things keep it honest. Compare against a trailing window rather than a fixed number, so the baseline moves with the business. Suppress alerts for known events — a retailer's seasonal catalogue shrink is not a defect. And separate severities: a 20% drop is a ticket, a 90% drop is a page.","The other discipline is to alert on the derivative, not the level. Nobody can say what the correct row count for a category is. Everybody can tell you that it should not have changed by a fifth overnight.",{"type":134,"title":135,"intro":136,"items":137},"list","What to record on every run, from day one","Cheap to write, impossible to reconstruct later.",[138,139,140,141,142,143],"URLs attempted and rows returned, as two separate numbers","HTTP status counts, broken out — 403s and 404s mean entirely different things","Per-field null counts","Which fallback rung answered, if you have a chain","Duration, which is the cheapest proxy for \"something changed\"","A content hash per row, so a value that never changes becomes visible",{"type":123,"variant":145,"title":146,"text":147,"cta":148},"product","Seeing this on your own feed","Scrapewise records attempted, returned, status breakdown and per-field coverage on every run, so the row-count delta and null-rate checks are readable without you building the plumbing. The judgement about what the number should have been is still yours — no vendor can know your catalogue.",{"label":149,"url":150},"Create a free account","https://portal.scrapewise.ai/register",{"type":85,"title":152,"intro":153,"headers":154,"rows":158},"Worked example: the same run, two success rates","One site, one night, one set of numbers. The dashboard reported 100% and the honest figure is about 74%, because the two calculations divide by different things. Nothing here is disputed — both rates are computed from the same run.",[155,156,157],"Measure","Count","Rate",[159,163,166,169,171,175],[160,161,162],"URLs on the link list","2,667","the denominator that matters",[164,161,165],"Pages attempted","",[167,168,165],"Pages fetched without an error","1,964",[170,168,165],"Rows written",[172,173,174],"Rows carrying a price, over rows written","1,964 of 1,964","100% — the figure on the dashboard",[176,177,178],"Rows carrying a price, over URLs attempted","1,964 of 2,667","about 74% — the figure to alert on",{"type":134,"title":180,"intro":181,"items":182},"What usually goes wrong","Monitoring that measures the job rather than the data produces a green dashboard above a feed nobody should be using.",[183,184,185,186,187,188],"Computing a success rate over rows returned. Every row returned has a price by definition, so the rate sits near 100% forever and the check is decorative.","Alerting on the level rather than the movement. Catalogues grow and shrink, so a fixed floor either fires every week or never fires; a percentage move against the trailing average tracks what is actually happening.","Giving every check the same severity. A doubling in the price null rate and one page returning 404 should not open the same ticket, or the ticket stops being read.","Monitoring across all sites at once. The one competitor that disappeared is a few per cent of the total, and a few per cent is indistinguishable from noise.","Having no hand check. Twenty rows a month compared against the live pages is the one control in the whole system that cannot itself drift.","Not recording attempted counts. Without them the honest denominator above cannot be computed at all, and the dashboard is permanently flattering.",{"type":80,"title":190,"paragraphs":191},"The reconciliation nobody does",[192,193,194],"Once a month, take twenty rows at random and open the pages by hand. Compare what the feed says to what the page says.","It takes half an hour and it is the only check that catches the category of error where everything is internally consistent and externally wrong — the right price from the wrong market, the list price where the promotion should be, yesterday's number served from a cache. Automated checks compare your data against your data. This is the only step that compares it against reality.","Write the result down with the date. Over a year it becomes the only honest answer you have to \"how much do we trust this feed\", and that question will be asked by someone senior at the worst possible moment.",[196,197,198,199],"The worst failures exit zero. Monitor output, not completion.","Row count versus trailing average is the highest-value single check — it catches most silent failures.","Success rate over returned rows is meaningless. The denominator must be what you attempted.","Alert on the derivative, separate ticket from page severity, and hand-check twenty rows a month against the live pages.",{"text":201,"label":202,"url":203},"Next: the decision itself — when to fix, when to work around, and when to stop.","Lesson 6: decide what to do when a site wins","/learn/keep-scrapers-alive/decide-what-to-do-when-a-site-wins",[205,211,217,222,227,228],{"slug":206,"navTitle":207,"title":208,"summary":209,"time":210,"needsAccount":73},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.","10 min",{"slug":212,"navTitle":213,"title":214,"summary":215,"time":216,"needsAccount":73},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.","12 min",{"slug":218,"navTitle":219,"title":220,"summary":221,"time":72,"needsAccount":73},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.",{"slug":223,"navTitle":224,"title":225,"summary":226,"time":72,"needsAccount":73},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":229,"navTitle":230,"title":231,"summary":232,"time":210,"needsAccount":73},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"slug":223,"navTitle":224,"title":225,"summary":226,"time":72,"needsAccount":73},{"slug":229,"navTitle":230,"title":231,"summary":232,"time":210,"needsAccount":73},1791047867282]