[{"data":1,"prerenderedAt":1038},["ShallowReactive",2],{"learn-course-keep-scrapers-alive":3,"learn-courses":826},{"slug":4,"order":5,"level":6,"time":7,"card_text":8,"seo":9,"hero":15,"outcomes":25,"who":34,"syllabus":45,"faq":48,"lessons":65},"keep-scrapers-alive",4,"You already have something running","6 lessons, about 65 minutes","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.",{"title":10,"description":11,"keywords":12,"og_title":13,"og_description":14},"Keep Scrapers Alive: A Free 6-Lesson Maintenance Course","Why web scrapers break and what to do about it. Reading failures properly, selectors that survive redesigns, bot walls, monitoring the feed rather than the run, and knowing when to stop. Free, ungated.","web scraper maintenance, scraper broke, scraper monitoring, css selector best practice, bot detection, scraper error handling, data quality monitoring, web scraping reliability","A free course on why scrapers break and how to keep them running","Six written lessons on scraper maintenance: failure diagnosis, durable selectors, bot walls, feed monitoring, and deciding when a site has won.",{"badge":16,"title":17,"subtitle":18,"cta_primary":19,"cta_secondary":22},"Course four","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.",{"label":20,"url":21},"Start with lesson one","/learn/keep-scrapers-alive/why-scrapers-break",{"label":23,"url":24},"See real run output","/custom-scrapers",{"title":26,"items":27},"What you will be able to do",[28,29,30,31,32,33],"Name the five reasons a scraper stops working, and tell them apart from the evidence rather than from a hunch","Read a failure down to its cause instead of restarting the job and hoping","Write selectors that survive a front-end rewrite, and recognise the ones that will not","Understand what a bot wall is actually measuring, and what genuinely changes the outcome","Monitor the output of a feed, not just whether the job exited zero","Decide, with a rule rather than a mood, when to stop fighting a site",{"title":35,"for_title":36,"for":37,"not_title":41,"not_for":42},"Who this is for","Written for",[38,39,40],"Anyone who already has scrapers running and keeps getting surprised by them","Developers who inherited somebody else's collection layer and want to stop firefighting it","Analysts who depend on a scraped feed and need to know how much to trust it","Not written for",[43,44],"First-time scraper authors — start with course one or course three, this assumes something already runs","Anyone looking for techniques to defeat a specific site's protection. That is not what this is, and lesson four explains why that framing loses.",{"title":46,"intro":47},"The six lessons","Read in order the first time. After that it works as a reference — lesson two is the one people come back to.",{"badge":49,"title":50,"description":51,"items":52},"FAQ","Before you start","What people usually want to know when a scraper has just broken.",[53,56,59,62],{"title":54,"description":55},"My scraper broke this morning. Which lesson do I read?","Lesson two. It is the diagnosis lesson and it is deliberately the longest. Most wasted maintenance time comes from fixing the wrong thing — rewriting a selector when the page never loaded, or rotating proxies when the product was simply discontinued.",{"title":57,"description":58},"Is this about avoiding detection?","No. Lesson four covers bot walls because you will meet them, but the framing there is about what the wall is measuring and what honestly changes the outcome, not about evasion. The most reliable answers turn out to be boring: slow down, be identifiable, and accept that some sites will say no.",{"title":60,"description":61},"Does any of this apply if I use a managed service?","The diagnosis and monitoring lessons especially. A vendor can own the fetching and the extraction, but nobody else can know that your category page should have returned two thousand rows and came back with four hundred. That judgement stays with you whatever you buy.",{"title":63,"description":64},"How much maintenance should I actually expect?","Measured across our own production groups, a storefront with no major redesign needs attention a handful of times a year, and a redesign costs a few hours per affected site. The trap is not the hours. It is that they arrive unannounced, all at once, and usually on the day something else is already on fire.",[66,190,307,446,569,697],{"slug":67,"nav_title":68,"title":69,"summary":70,"time":71,"needs_account":72,"seo":73,"blocks":77,"takeaways":181,"next_step":186},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.","10 min",false,{"title":74,"description":75,"keywords":76},"Why Web Scrapers Break: The Five Real Causes","Redesigns, bot walls, rendering changes, content drift and silent partial failure. What each one looks like, how often it happens, and why the fixes are different.","why scrapers break, web scraper maintenance, scraper stopped working, scraper reliability, data pipeline failure",[78,83,118,124,129,139,167,176],{"type":79,"paragraphs":80},"prose",[81,82],"A scraper that worked yesterday and does not work today has failed for one of five reasons. They look similar from the outside — an empty result, or fewer rows than usual — and they have almost nothing in common underneath. The habit worth building is naming the category before touching any code, because four of the five fixes are wrong four fifths of the time.","In rough order of how often they occur on a mature setup:",{"type":84,"title":85,"intro":86,"headers":87,"rows":92},"table","The five failure modes","Frequency here is from our own production groups, not an industry figure. Your mix will differ by sector; the categories will not.",[88,89,90,91],"Failure","What you see","What actually changed","Typical fix",[93,98,103,108,113],[94,95,96,97],"Content drift","Fewer rows, no errors","Products delisted, renamed or moved; the site is fine","Nothing. Update the URL list.",[99,100,101,102],"Redesign","Zero rows, page fetched fine","The markup moved; your selector points at nothing","Re-point the selector",[104,105,106,107],"Bot wall","Non-200 status, or a page that is not the product page","The site decided you are not a browser","Slow down, identify yourself, or route differently",[109,110,111,112],"Rendering change","Page fetched, markup present, values empty","Content moved behind JavaScript","Render the page, or read the data source directly",[114,115,116,117],"Silent partial","Looks normal, is wrong","Pagination capped, a region defaulted, a currency flipped","The hardest. Lesson five.",{"type":79,"title":119,"paragraphs":120},"The one that costs the most is not the one you expect",[121,122,123],"Redesigns feel like the enemy because they are dramatic — everything goes to zero and somebody notices within the hour. That visibility is a gift. A loud failure is a cheap failure.","The expensive category is the last row. A run that returns four hundred rows where it used to return two thousand exits successfully. The scheduler is happy. The dashboard has data in it. And every number downstream is now computed over a fifth of the catalogue, which is worse than having no number at all, because somebody will price against it.","We have measured exactly this on a live client group: rows per week fell from 10,382 to 8,226 across two weeks while the total number of requests stayed identical — 11,032 then 11,034. Nothing errored. Nothing alerted. The job had been fetching and billing for every page the whole time and simply returning less.",{"type":125,"variant":126,"title":127,"text":128},"callout","warning","Exit code zero is not a health check","The single most common monitoring mistake is treating \"the job finished\" as \"the job worked\". Those are different claims and only one of them is worth an alert. Lesson five is entirely about the difference, and it is the lesson that changes the most for most teams.",{"type":130,"title":131,"intro":132,"items":133},"list","Questions that place a failure in the right category in under a minute","Ask these in order. The first one that gives an interesting answer is usually the whole diagnosis.",[134,135,136,137,138],"What HTTP status came back? A non-200 is a bot wall or an outage, and no selector work will help.","Did the page body arrive, and is it the page you asked for? A 200 that returns a challenge page is still a block.","Is the field present in the markup but empty, or absent entirely? Present-but-empty points at rendering; absent points at a redesign.","Did every URL fail, or a subset? A subset is almost always content drift.","Did the row count fall, or go to zero? Zero is loud and simple. A fall is the dangerous one.",{"type":84,"title":140,"intro":141,"headers":142,"rows":146},"Worked example: one week of symptoms, placed in five minutes","Five scrapers, five symptoms, five different fixes. The column to read is the middle one: the evidence that separates the categories is cheap to collect, and almost nobody collects it, which is why maintenance feels like one endless problem rather than five bounded ones.",[143,144,145],"Symptom","The evidence that settles it","Category, and what the fix is",[147,151,155,159,163],[148,149,150],"Zero rows, job red, connection refused","HTTP status on a single URL","Bot wall — slow down, or stop and document the gap",[152,153,154],"Zero rows, job green, 200 returned","Is the price in the HTML source, or only on screen?","Rendering change — the page went client-side",[156,157,158],"Price null on every row, everything else intact","Search the response body for the literal price string","Content drift — one selector, five minutes",[160,161,162],"Every field null, 200 returned, body unrecognisable","Open the page in a browser","Redesign — rewrite the extraction for this site",[164,165,166],"Rows down 21% over two weeks, no errors anywhere","Row count against the trailing average, per site","Silent partial — the expensive one",{"type":130,"title":168,"intro":169,"items":170},"What usually goes wrong","Four of these five are attempts to fix a scraper before anyone has said which of the five things is wrong with it.",[171,172,173,174,175],"Treating all five as one problem called maintenance. They have four different fixes and one of them is \"stop\", so the category has to be named before anyone touches code.","Reaching for proxies first. It is the standard answer to one of the five modes, the wrong answer to the other four, and expensive in all five.","Using exit code zero as the health signal. The worst mode in the table above returns successfully, every day, for a fortnight.","Diagnosing from a browser. Yours carries cookies, history and a residential address; the job has none of them, which is why the page you are looking at is not the page it received.","Rewriting the selector before checking whether the value is in the source at all. If the price only exists after JavaScript runs, no selector will ever find it, and you can spend a day proving that.",{"type":79,"title":177,"paragraphs":178},"Why it matters that these are five and not one",[179,180],"Because the reflex fix for each is useless for the other four. Rotating proxies does nothing about a redesign. Rewriting a selector does nothing about a product that no longer exists. Adding a browser renderer to a site that was never JavaScript-driven costs you money per page and fixes nothing.","Most teams that describe scraping as \"constant maintenance\" are not doing more maintenance than anyone else. They are doing the same amount, spent on the wrong category, which means the actual cause survives the fix and comes back next week.",[182,183,184,185],"Five failure modes: content drift, redesign, bot wall, rendering change, silent partial.","Loud failures are cheap. The expensive one returns successfully with less data.","A measured case: rows fell 21% across two weeks with request volume unchanged and zero errors.","Name the category before you touch code — four of the five standard fixes are wrong most of the time.",{"text":187,"label":188,"url":189},"Next: how to actually read a failure instead of guessing which of the five it was.","Lesson 2: read the failure, not the symptom","/learn/keep-scrapers-alive/read-the-failure-not-the-symptom",{"slug":191,"nav_title":192,"title":193,"summary":194,"time":195,"needs_account":72,"seo":196,"blocks":200,"takeaways":298,"next_step":303},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.","12 min",{"title":197,"description":198,"keywords":199},"How to Debug a Broken Web Scraper: A Diagnosis Routine","A repeatable routine for diagnosing scraper failures: status before body, body before selector, one URL before the whole run. Includes the three false conclusions that waste the most time.","debug web scraper, scraper error, scraper returns nothing, scraper troubleshooting, http 403 scraper, scraper debugging",[201,205,228,232,238,265,283,292],{"type":79,"paragraphs":202},[203,204],"The instinct when a run comes back empty is to open the page in a browser, see the price sitting there in plain sight, and conclude that the selector is broken. That conclusion is right perhaps a third of the time, and the other two thirds are expensive, because you will rewrite a selector that was never wrong and the run will stay broken.","The routine below is ordered so that each step can only be reached if the previous one has been eliminated. It is deliberately mechanical. The whole point is to remove the guessing.",{"type":206,"title":207,"intro":208,"items":209},"steps","The routine","Do these in order on a single URL. Never debug against the whole run — you cannot see anything in five thousand rows of log.",[210,213,216,219,222,225],{"title":211,"text":212},"Pick one failing URL and work only on that","One URL that definitely used to work. If you cannot name one, that is itself the finding: you may be looking at content drift rather than breakage.",{"title":214,"text":215},"Check the status code before anything else","A 403, 429 or 503 ends the investigation. That is a bot wall or a rate limit, and nothing in your extraction layer is involved. Jump to lesson four.",{"title":217,"text":218},"Check what the body actually is","A 200 is not proof you got the product page. Challenge pages, consent interstitials, geo-redirects and soft 404s all return 200 with a perfectly valid HTML body. Look at the title tag and the length. A product page that is suddenly 4 KB is not a product page.",{"title":220,"text":221},"Search the raw body for the value, not for your selector","Take the price you can see in the browser and grep the fetched body for the digits. If the number is in there, you have an extraction problem. If it is not, the page you fetched is not the page you looked at, and that is a rendering or routing problem.",{"title":223,"text":224},"Only now look at the selector","You have earned the right to. And at this point the fix is usually obvious, because you know the value is present and you know where.",{"title":226,"text":227},"Re-run the single URL, then ten, then the group","Fixing one and immediately launching five thousand is how a wrong fix becomes an expensive wrong fix.",{"type":125,"variant":229,"title":230,"text":231},"note","Grep the body before you theorise","Searching the fetched HTML for the literal digits of the price takes ten seconds and splits the problem space in half every time. It is the single highest-value habit in this course. We have skipped it, shipped a wrong fix on the strength of a plausible theory, and had to ship a second fix to undo the first.",{"type":79,"title":233,"paragraphs":234},"Three false conclusions this routine is built to prevent",[235,236,237],"The first is \"it works in my browser, so the site is fine\". Your browser has cookies, a residential IP, a full JavaScript engine and a history with that domain. It is not a control group. A page that loads for you and refuses a datacenter address is the ordinary case, not an anomaly — and it means a failure you cannot reproduce locally is still real.","The second is \"it returned 200, so we got the page\". Covered above, and it is the one that burns the most hours, because a 200 feels conclusive.","The third is subtler: \"the fix worked, the run is green\". Green after a fix means the run completed. It does not mean the rows are right. Compare the row count to last week's before you call it done — a selector that now matches a different element will happily return five thousand rows of the wrong thing.",{"type":84,"title":239,"headers":240,"rows":244},"Symptom to cause, once you have the evidence",[241,242,243],"Evidence","Cause","Where to go",[245,249,252,256,258,261],[246,247,248],"403 / 429 / 503","Bot wall or rate limit","Lesson 4",[250,251,248],"200, tiny body, unexpected title","Challenge or consent interstitial",[253,254,255],"200, full body, value absent from source","Rendered client-side","Lesson 3",[257,99,255],"200, value present in body, selector misses",[259,94,260],"Some URLs fine, some 404","Refresh the URL list",[262,263,264],"Everything green, fewer rows","Silent partial failure","Lesson 5",{"type":206,"title":266,"intro":267,"items":268},"Worked example: ten minutes on one URL","A feed that wrote 1,964 rows from 2,667 attempted pages. The temptation is to theorise — delisted products, a bad selector, a blocked range — and every theory sounds plausible. The routine answers it without a theory, and no step costs more than two minutes.",[269,272,275,278,280],{"title":270,"text":271},"Check the status distribution across the whole run, not a sample","If the failures all carry one status class you have the answer in ninety seconds. In this case every failure was a page-level error and not one was an empty result, which already kills the delisted-products theory: a delisted page returns something, it does not fail.",{"title":273,"text":274},"Fetch one failing URL and read the status","A 403 is a bot wall. A 404 is a stale URL list. A 200 means the cause is further down the page and the next two steps are worth the time.",{"title":276,"text":277},"Search the raw body for the literal value","Before looking at any selector, search the response for the price exactly as it appears on screen. Present means a selector problem. Absent means a rendering problem, and no amount of selector work will ever fix it.",{"title":223,"text":279},"Three of the four candidate causes have been eliminated for the price of four fetches. This is the point at which editing the selector is an informed act rather than a guess.",{"title":281,"text":282},"Prove the path ran before explaining why it failed","If a fix was deployed and nothing changed, establish that the new code executed at all — a timing difference, a log line, a billing delta at the provider. Explaining the failure of a code path that never ran is the most expensive way to spend an afternoon that is available to anybody.",{"type":130,"title":168,"intro":284,"items":285},"Each of these turns a ten-minute diagnosis into a two-day one.",[286,287,288,289,290,291],"Theorising before fetching. Delisted products, a supplier feed change and a broken selector all explain the same row count, and a single fetch separates them.","Reading the status and stopping there. A 200 routinely carries a consent interstitial, a geo-redirect or a challenge page, all of which parse to nothing at all.","Testing from your own machine and concluding the site is fine. Your laptop has cookies, history and a residential address; the job has a flagged data centre range. The site treats them as different visitors, because they are.","Sampling the failures. The status distribution across the whole run is one query and it frequently ends the investigation by itself.","Changing three things and redeploying. When it works nobody knows which change worked, and when it does not nobody knows which one to undo.","Explaining a failure in code that never executed. Prove it ran — timing, logs, or a billing delta — before explaining anything about it.",{"type":79,"title":293,"paragraphs":294},"Prove the path ran before you explain why it failed",[295,296,297],"This is the rule that took us longest to learn and it generalises past scraping. Before building any theory about why a step failed, establish that the step executed at all.","Timing is the cheapest evidence: a fetch that supposedly hit a renderer and returned in 120 milliseconds did not hit a renderer. Billing is the next cheapest, if your provider exposes a credit balance — make one known-good call, note the delta, then diff the balance around the request you are investigating. If the counter did not move, the code path you are theorising about never ran, and every theory about its behaviour is noise.","We shipped one wrong fix and one wrong explanation in a single session by skipping that check. Both were plausible. Neither touched the actual cause.",[299,300,301,302],"Status, then body, then value-in-source, then selector. In that order, on one URL.","A 200 does not mean you got the page you asked for.","Your browser is not a control group — it has cookies, a residential IP and history.","Prove the code path executed before explaining why it misbehaved. Timing and billing deltas are the cheapest proof.",{"text":304,"label":305,"url":306},"Next: the extraction side — which selectors survive a rewrite and which are built to fail.","Lesson 3: selectors that survive a redesign","/learn/keep-scrapers-alive/selectors-that-survive-a-redesign",{"slug":308,"nav_title":309,"title":310,"summary":311,"time":312,"needs_account":72,"seo":313,"blocks":317,"takeaways":437,"next_step":442},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.","11 min",{"title":314,"description":315,"keywords":316},"CSS Selectors That Survive a Website Redesign","Which extraction targets are durable and which break on the next deploy. Structured data, test IDs, semantic markup and generated class names ranked, plus how to build a fallback chain.","css selector best practice, xpath vs css selector, scraper selector broke, json-ld scraping, structured data extraction, durable web scraping",[318,322,371,374,379,388,421,430],{"type":79,"paragraphs":319},[320,321],"Two scrapers can read the same price off the same page and one of them will still be working in two years. The difference is not skill at writing selectors. It is what they chose to point at.","Every extraction target has an implicit contract with the site's front-end team, and most of those contracts are imaginary. A generated class name is not a promise. A deeply nested div path is not a promise. Here is what is ranked by how real the promise is.",{"type":84,"title":323,"intro":324,"headers":325,"rows":330},"Extraction targets, most durable first","\"Survives\" means it kept working across a front-end rewrite, not just a CSS tweak.",[326,327,328,329],"Target","Durability","Why","Catch",[331,336,341,346,351,356,361,366],[332,333,334,335],"JSON-LD / structured data","Very high","It exists to be machine-read and the SEO team defends it","Not every site publishes it, and some publish it wrong",[337,338,339,340],"data-testid / data-qa attributes","High","Their own test suite breaks if it changes","Not public API — can vanish in a tooling migration",[342,343,344,345],"Microdata (itemprop)","High where present","Same incentive as JSON-LD","Increasingly rare; measured on 60 major retailers, zero used it",[347,348,349,350],"Semantic HTML and ARIA roles","Medium-high","Accessibility work tends to survive restyling","Inconsistently applied",[352,353,354,355],"Stable id attributes","Medium","Often tied to backend templates","Framework rewrites take them out",[357,358,359,360],"Human-readable class names","Low-medium","Meaningful to a developer, so sometimes preserved","No guarantee whatsoever",[362,363,364,365],"Generated class names (css-1x2y3z)","Very low","They are a build artefact","Change on a dependency bump, with no visible change to the page",[367,368,369,370],"Positional paths (div > div:nth-child(3))","Lowest","They encode layout, and layout is what changes","Breaks when someone adds a banner",{"type":125,"variant":229,"title":372,"text":373},"Check for structured data before you write anything","Fetch the page and search the source for application/ld+json. If a Product block is there with a price in it, you are done, and that extraction will likely outlive three redesigns. It takes thirty seconds to check and people routinely skip it and hand-write a selector against markup that will be gone by spring.",{"type":79,"title":375,"paragraphs":376},"The honest ceiling on structured data",[377,378],"It is also worth knowing how far it gets you, because \"just use the schema\" is advice people give without having measured it. Across thirteen real storefront fixtures we tested, eleven published nothing machine-readable at all on the product page. Of those that did, several published a JSON-LD block whose price disagreed with the price rendered on screen, usually because it was the list price and the page was showing a promotion.","So: always check, often you will be lucky, and never assume the structured value is the one a customer sees. Cross-check one page by hand before you trust a thousand.",{"type":130,"title":380,"items":381},"Rules that make a selector last",[382,383,384,385,386,387],"Anchor to meaning, not position. Something that describes what the element is will outlive something that describes where it sits.","Shorter is stronger. Each additional step in a path is another thing that can move.","Never match on a class that looks generated. If it contains a hash, it is a build artefact.","Prefer an attribute the site's own tests depend on. Their CI is now protecting your scraper for free.","Extract the raw value and convert later. Pull the number and the currency as they appear, and do unit conversion as a separate step — then a formatting change on the site does not become a parsing bug.","Write down what you expect. A note saying \"price is in the JSON-LD offers block, fallback is the h1 sibling\" turns a future outage into a five-minute fix for whoever is on call.",{"type":84,"title":389,"intro":390,"headers":391,"rows":396},"Worked example: a chain that reports its own decay","Four rungs on one site, with the share of pages each one answered in January and again in June. The chain never failed and nothing ever alerted. What changed is that the feed moved off a source somebody defends and onto one nobody does — which is precisely the state the next redesign will break. It is visible only because each run recorded which rung produced the answer.",[392,393,394,395],"Rung","Source","Share in January","Share in June",[397,402,407,412,417],[398,399,400,401],"1","JSON-LD Product, offers.price","62%","0%",[403,404,405,406],"2","A data-testid attribute","31%","58%",[408,409,410,411],"3","A CSS class selector","6%","39%",[413,414,415,416],"4","Regex over the visible text","1%","3%",[418,419,420,420],"","Pages that produced a price","100%",{"type":130,"title":168,"intro":422,"items":423},"A selector is written once and then lives for years, so the failure modes are all about time rather than correctness.",[424,425,426,427,428,429],"Writing a selector against a generated class name. Those strings are build output: nobody is defending them and they change without any person deciding to change them.","Assuming structured data will be there. On one measured set of thirteen storefront fixtures, eleven published nothing machine-readable on the product page. Check, rather than designing around the assumption.","Trusting structured data that is there without comparing it to the page. It is the most stable source available and also the one most often left stale after a promotion.","Building a silent fallback chain. It converts a loud failure into a gradual decay, and you discover it when the last rung goes as well.","Taking the first match. A page can carry the same price string in a header, a schema block and a recently-viewed carousel; counting votes across candidates is more robust than taking whichever came first.","Selecting on position. Third div inside the second section works perfectly until somebody adds a banner, which is a thing marketing does without telling engineering.",{"type":79,"title":431,"paragraphs":432},"Build a fallback chain, but make it noisy",[433,434,435,436],"The strong pattern is an ordered chain: try the structured data, then the test attribute, then the semantic element, then the hand-written selector. First one that yields a plausible value wins.","The pattern has one failure mode and it is severe. A silent chain hides decay. If rung one stopped working in March and rung four has been quietly covering for it ever since, you will find out in September when rung four breaks too — and by then nobody remembers what rung one was for.","So record which rung answered. Not an alert, just a field on the row. When the distribution shifts — eighty percent of rows suddenly answering from rung three instead of rung one — that is the redesign, caught weeks before it would have surfaced as an outage.","One more detail from a bug that reached production: when a chain votes, count the votes rather than taking the first match. A page can contain the same attribute fifty times, with forty-nine of them empty stubs. First-match returns the stub. So does last-match. Counting what the rungs actually agree on returns the price.",[438,439,440,441],"Rank targets by whose incentive protects them: structured data and test IDs have defenders, generated class names do not.","Check for a JSON-LD Product block before writing any selector — but expect most sites not to have one.","Measured: 11 of 13 storefront fixtures published nothing machine-readable on the product page.","Fallback chains should record which rung answered, or they hide decay until every rung is gone.",{"text":443,"label":444,"url":445},"Next: what a bot wall is actually measuring, and what honestly changes the outcome.","Lesson 4: bot walls and what actually works","/learn/keep-scrapers-alive/bot-walls-and-what-actually-works",{"slug":447,"nav_title":448,"title":449,"summary":450,"time":312,"needs_account":72,"seo":451,"blocks":455,"takeaways":560,"next_step":565},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"title":452,"description":453,"keywords":454},"Why Scrapers Get Blocked: Bot Detection Explained","What bot detection actually measures, why a flagged datacenter IP is refused rather than challenged, and which responses genuinely change the outcome. Plus when to stop.","bot detection, scraper blocked 403, cloudflare scraper, rate limiting scraper, datacenter ip blocked, residential proxy, web scraping blocked",[456,460,487,490,498,503,512,546,555],{"type":79,"paragraphs":457},[458,459],"A protection layer is not trying to work out whether you are a robot. It is scoring how much you cost and how much you are worth, and the score is mostly made of three things: where you are connecting from, how fast you are asking, and whether your request looks like it came from the software it claims to come from.","That framing matters, because it tells you which of your options are real. You cannot argue with the score. You can change the inputs to it, and only some of those inputs are yours to change.",{"type":84,"title":461,"headers":462,"rows":466},"What gets measured, and what you can do about it",[463,464,465],"Signal","What it sees","Your realistic lever",[467,471,475,479,483],[468,469,470],"IP reputation","Datacenter range, known provider, prior abuse from that block","Route differently, or accept the refusal",[472,473,474],"Request rate","Requests per minute from one source, and the regularity of the gaps","Slow down and vary. This is the biggest free win.",[476,477,478],"Header coherence","Whether your headers describe a real client consistently","Send a coherent, honest set. Do not impersonate.",[480,481,482],"Session behaviour","Straight to product pages, no cookies, no referrer, never a listing","Behave like a reader: accept cookies, follow links",[484,485,486],"Browser fingerprint","JavaScript challenges, canvas, timing","Render properly, or stop",{"type":125,"variant":126,"title":488,"text":489},"The laptop test will lie to you","A flagged datacenter address is often refused outright, not challenged. So the body fingerprint you carefully measured from your own machine — the challenge page, the specific wording, the script tag you keyed your detection on — never appears on the deployment host, which simply gets a 403. Any detection logic built from a local reproduction is a false negative in production. Key your gate on the HTTP status, which is the one signal that survives the move.",{"type":79,"title":491,"paragraphs":492},"The boring answers outperform the clever ones",[493,494,495,496,497],"Nearly everything written about this topic is about defeating detection, and nearly all of it ages badly, because it describes a specific countermeasure against a specific version of a specific vendor's product. The things that keep working are unglamorous.","Ask less often. Most scraping is dramatically over-frequent relative to how often the underlying data changes. A price that moves twice a week does not need hourly polling, and halving your request rate removes the signal you were most visibly failing on.","Be identifiable. A user agent that says who you are and a contact address is not naivety — it is the thing that gets you an allowlist instead of a ban when somebody looks at the logs. Pretending to be Chrome and failing the fingerprint check is worse than both.","Ask for less. Pulling a thousand product pages when the category listing already carries the price and the title is a thousand requests you did not need, and a hundredfold increase in your visibility.","Cache aggressively. The cheapest request is the one you did not make.",{"type":79,"title":499,"paragraphs":500},"One measured thing about market and locale",[501,502],"Related, and it has cost us real money. Large international retailers decide which market you are in, and that decision changes the price you are served — not the formatting, the actual number.","On Amazon we measured this directly. Sending an Accept-Language header alone did nothing at all: three attempts within the same minute, same market, same price every time. The cookie is what flips the market. Get that wrong and your scraper does not fail — it succeeds, and returns a correct price from the wrong country, which is the silent-partial failure from lesson one wearing a different coat. A single-currency sanity check on the output would have caught it in a day. We did not have one, and the wrong number was live for longer than it should have been.",{"type":130,"title":504,"intro":505,"items":506},"When to stop","Some sites will not be read, and continuing is a cost with no return. Stop when any two of these are true.",[507,508,509,510,511],"You have spent more than a day on one site and the success rate is still under half","The fix involves solving a visual or interactive challenge","You are being asked to log in, and the terms attached to that login forbid automated access","The per-page cost of getting through exceeds what the data is worth to you","The thing you need is published somewhere else — a feed, a marketplace listing, an aggregator — that is not fighting you",{"type":84,"title":513,"intro":514,"headers":515,"rows":519},"Worked example: the escalation ladder, priced","Each rung costs more than the one above it, and the order matters because the cheap rungs resolve most cases. Cost is given as a multiple of a plain fetch rather than in currency, because the absolute figures differ by provider and by month while the ratios are the part that generalises — and the ratios are what should govern the decision.",[392,516,517,518],"What it changes","Relative cost","When it is the right answer",[520,524,527,532,536,541],[521,472,522,523],"Slow down","×1","Almost always first — rate is the cheapest thing you control",[525,476,522,526],"Fix the headers","A client whose headers match no real browser",[528,529,530,531],"Render the page","Execution of page scripts","about ×5","The value is absent from the source and present on screen",[533,468,534,535],"Residential address","about ×10","A data centre range is being refused outright",[537,538,539,540],"Both together","Both","about ×25","Rarely, and never without a spend cap in front of it",[542,543,544,545],"Stop","Nothing","×0","A documented coverage gap — the next lesson is about writing it",{"type":130,"title":168,"intro":547,"items":548},"Almost every expensive mistake here is an escalation made before the cheap rungs were tried, or measured somewhere the job does not run.",[549,550,551,552,553,554],"Starting at the expensive rung. Residential addresses and full rendering fix a minority of cases and cost five to twenty-five times as much on every page, permanently.","Reproducing it on a laptop. A flagged data centre address is refused outright rather than challenged, so a body fingerprint measured from your machine is a false negative on the host that actually runs the job. Key the detection on the HTTP status.","Escalating without a spend cap. The combined rung is roughly twenty-five times the base rate, and a retry loop on top of it is how a month's budget disappears in an afternoon.","Changing the market without meaning to. On at least one large retailer the cookie decides the market and the language header alone does nothing, so an escalation can return a perfectly correct price from the wrong country.","Impersonating a browser as a strategy. It is an arms race against a team whose whole job it is, and it converts a conduct question into a bad-faith one — course six covers what that costs.","Treating stopping as a failure. A documented coverage gap with numbers attached is a result; a partial feed nobody labelled is a liability.",{"type":79,"title":556,"paragraphs":557},"Stopping is a result",[558,559],"A coverage note saying \"this retailer is not readable, here is the evidence, here is what we propose instead\" is a legitimate deliverable. It is worth considerably more than a feed that is quietly thirty percent wrong because somebody refused to be beaten.","The failure mode to avoid is the one where a site that mostly blocks you produces a trickle of rows and nobody labels them. A partial feed presented as a complete one is the most expensive artefact in this whole discipline.",[561,562,563,564],"Detection scores IP reputation, request rate, header coherence and session behaviour. Rate is the one you control most cheaply.","A flagged datacenter IP is refused, not challenged — detection logic built from a laptop reproduction is a false negative in production. Key on HTTP status.","Measured on Amazon: the cookie flips the market, Accept-Language alone does nothing. Wrong market means a correct price from the wrong country.","Stopping with a documented coverage gap beats a partial feed nobody labelled.",{"text":566,"label":567,"url":568},"Next: the failure that does not announce itself — monitoring output rather than exit codes.","Lesson 5: monitor the feed, not the run","/learn/keep-scrapers-alive/monitor-the-feed-not-the-run",{"slug":570,"nav_title":571,"title":572,"summary":573,"time":312,"needs_account":72,"seo":574,"blocks":578,"takeaways":688,"next_step":693},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"title":575,"description":576,"keywords":577},"Data Quality Monitoring for Web Scrapers: Six Checks","How to detect a scraper that completes successfully but returns wrong data. Row count deltas, null rates, distribution shifts, staleness, and thresholds that avoid alert fatigue.","data quality monitoring, scraper monitoring, data pipeline alerting, silent data failure, row count anomaly detection, price data validation",[579,583,615,620,623,629,639,646,673,682],{"type":79,"paragraphs":580},[581,582],"Everything in this lesson exists because of one fact: the worst scraper failures do not raise errors. They complete, they write rows, the dashboard refreshes, and the numbers are wrong.","Monitoring the run tells you the job finished. Monitoring the feed tells you whether the thing it produced is usable. Only the second one is worth waking up for, and most teams have only built the first.",{"type":84,"title":584,"intro":585,"headers":586,"rows":590},"Six checks, cheapest first","Implement them in this order. The first two cover most of the real-world damage.",[587,588,589],"Check","Catches","Suggested trigger",[591,595,599,603,607,611],[592,593,594],"Row count versus trailing average","Pagination caps, partial blocks, truncated lists","Deviation beyond about 20% of the 7-day mean",[596,597,598],"Null rate per field","A single selector breaking while the rest hold","Any field whose null rate doubles week on week",[600,601,602],"Value distribution","Currency flips, unit errors, wrong market","Median moves more than 15% with no known cause",[604,605,606],"Staleness","A cached response being re-served as fresh","Any row whose value is byte-identical for an implausible stretch",[608,609,610],"Coverage against the expected URL list","Silent drops you never asked about","Attempted versus returned falls below your agreed floor",[612,613,614],"Cross-source agreement","Everything else, where you have a second source","Two sources disagreeing by more than a tolerance",{"type":79,"title":616,"paragraphs":617},"Row count is the single highest-value check",[618,619],"If you only ever build one of these, build this one. Store the count per source per run and compare it to the trailing average. It is a handful of lines and it catches the majority of silent failures, because almost every silent failure shows up first as \"less than usual\".","The live case from lesson one is exactly this shape. Rows per week went 10,382 then 8,237 then 8,226, while total requests stayed at 11,032, 11,035 and 11,034. Every page was fetched. Every page was billed. A fifth of the output simply stopped arriving, and because the error count was zero and the job was green, nothing anywhere said so. A row-count delta check would have flagged it in week one.",{"type":125,"variant":126,"title":621,"text":622},"Count what was attempted, not just what succeeded","A success rate computed over the rows you got is meaningless — it is always close to 100%, by construction. The denominator has to be what you intended to fetch. Attempted versus returned is the number that tells you the truth, and it is the number that is hardest to get after the fact, which is why it has to be recorded at run time.",{"type":79,"title":624,"paragraphs":625},"Thresholds, and the alert nobody reads",[626,627,628],"A monitor that fires every day is not a monitor, it is a background noise generator, and the second week of it is more dangerous than having no monitor at all because now everyone has learned to dismiss that channel.","Three things keep it honest. Compare against a trailing window rather than a fixed number, so the baseline moves with the business. Suppress alerts for known events — a retailer's seasonal catalogue shrink is not a defect. And separate severities: a 20% drop is a ticket, a 90% drop is a page.","The other discipline is to alert on the derivative, not the level. Nobody can say what the correct row count for a category is. Everybody can tell you that it should not have changed by a fifth overnight.",{"type":130,"title":630,"intro":631,"items":632},"What to record on every run, from day one","Cheap to write, impossible to reconstruct later.",[633,634,635,636,637,638],"URLs attempted and rows returned, as two separate numbers","HTTP status counts, broken out — 403s and 404s mean entirely different things","Per-field null counts","Which fallback rung answered, if you have a chain","Duration, which is the cheapest proxy for \"something changed\"","A content hash per row, so a value that never changes becomes visible",{"type":125,"variant":640,"title":641,"text":642,"cta":643},"product","Seeing this on your own feed","Scrapewise records attempted, returned, status breakdown and per-field coverage on every run, so the row-count delta and null-rate checks are readable without you building the plumbing. The judgement about what the number should have been is still yours — no vendor can know your catalogue.",{"label":644,"url":645},"Create a free account","https://portal.scrapewise.ai/register",{"type":84,"title":647,"intro":648,"headers":649,"rows":653},"Worked example: the same run, two success rates","One site, one night, one set of numbers. The dashboard reported 100% and the honest figure is about 74%, because the two calculations divide by different things. Nothing here is disputed — both rates are computed from the same run.",[650,651,652],"Measure","Count","Rate",[654,658,660,663,665,669],[655,656,657],"URLs on the link list","2,667","the denominator that matters",[659,656,418],"Pages attempted",[661,662,418],"Pages fetched without an error","1,964",[664,662,418],"Rows written",[666,667,668],"Rows carrying a price, over rows written","1,964 of 1,964","100% — the figure on the dashboard",[670,671,672],"Rows carrying a price, over URLs attempted","1,964 of 2,667","about 74% — the figure to alert on",{"type":130,"title":168,"intro":674,"items":675},"Monitoring that measures the job rather than the data produces a green dashboard above a feed nobody should be using.",[676,677,678,679,680,681],"Computing a success rate over rows returned. Every row returned has a price by definition, so the rate sits near 100% forever and the check is decorative.","Alerting on the level rather than the movement. Catalogues grow and shrink, so a fixed floor either fires every week or never fires; a percentage move against the trailing average tracks what is actually happening.","Giving every check the same severity. A doubling in the price null rate and one page returning 404 should not open the same ticket, or the ticket stops being read.","Monitoring across all sites at once. The one competitor that disappeared is a few per cent of the total, and a few per cent is indistinguishable from noise.","Having no hand check. Twenty rows a month compared against the live pages is the one control in the whole system that cannot itself drift.","Not recording attempted counts. Without them the honest denominator above cannot be computed at all, and the dashboard is permanently flattering.",{"type":79,"title":683,"paragraphs":684},"The reconciliation nobody does",[685,686,687],"Once a month, take twenty rows at random and open the pages by hand. Compare what the feed says to what the page says.","It takes half an hour and it is the only check that catches the category of error where everything is internally consistent and externally wrong — the right price from the wrong market, the list price where the promotion should be, yesterday's number served from a cache. Automated checks compare your data against your data. This is the only step that compares it against reality.","Write the result down with the date. Over a year it becomes the only honest answer you have to \"how much do we trust this feed\", and that question will be asked by someone senior at the worst possible moment.",[689,690,691,692],"The worst failures exit zero. Monitor output, not completion.","Row count versus trailing average is the highest-value single check — it catches most silent failures.","Success rate over returned rows is meaningless. The denominator must be what you attempted.","Alert on the derivative, separate ticket from page severity, and hand-check twenty rows a month against the live pages.",{"text":694,"label":695,"url":696},"Next: the decision itself — when to fix, when to work around, and when to stop.","Lesson 6: decide what to do when a site wins","/learn/keep-scrapers-alive/decide-what-to-do-when-a-site-wins",{"slug":698,"nav_title":699,"title":700,"summary":701,"time":71,"needs_account":72,"seo":702,"blocks":706,"takeaways":817,"next_step":822},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"title":703,"description":704,"keywords":705},"When to Stop Trying to Scrape a Website","A decision rule for repairing, routing around, or abandoning a source. How to value the data, cap the effort, and write a coverage gap report that is useful.","when to stop scraping, scraping coverage gap, data sourcing decision, alternative data source, scraper cost benefit",[707,711,729,735,754,757,764,799,808],{"type":79,"paragraphs":708},[709,710],"At some point a source becomes more expensive than it is worth, and the discipline is to notice that on purpose rather than by exhaustion. Teams that handle this badly do not usually make the wrong call — they make no call at all, and spend six months half-maintaining something nobody has decided to keep.","There are three available answers and it is worth being explicit about which one you are choosing.",{"type":84,"title":712,"headers":713,"rows":717},"The three answers",[714,715,716],"Choice","When it is right","What it costs",[718,722,726],[719,720,721],"Repair","Cause is identified, fix is bounded, source is load-bearing","Hours now, and the same hours again at the next redesign",[723,724,725],"Route around","The same data exists somewhere less defended","A different source means different coverage — say so",[542,727,728],"Cost per page exceeds the value, or the block is categorical","A documented gap, which is a real cost, honestly priced",{"type":79,"title":730,"paragraphs":731},"Put a number on the data before you argue about the effort",[732,733,734],"Most of these debates go badly because one side is talking about difficulty and the other about importance, and neither has been quantified.","The question that resolves it: if this source disappeared tomorrow, what decision gets made worse? If the answer is \"our weekly price position on four hundred SKUs we actively compete on\", that is load-bearing and you should spend real effort. If the answer is \"a column on a dashboard that two people open\", you have your decision and it is not the one anyone was arguing for.","Then cap the effort before starting. A day, a week, whatever is proportionate — decided in advance, because the sunk cost will absolutely argue for one more afternoon and it will say that every afternoon.",{"type":206,"title":736,"intro":737,"items":738},"The routing-around checklist","Before concluding a source is unavailable, the same data is often published somewhere nobody thought to look.",[739,742,745,748,751],{"title":740,"text":741},"The site's own feed","Sitemaps, product feeds, affiliate exports and RSS are published deliberately and defended far less than the HTML.",{"title":743,"text":744},"A marketplace listing","Plenty of retailers that block you directly also sell on a marketplace that does not, with the same prices attached.",{"title":746,"text":747},"A comparison site or aggregator","Lower precision and a lag, but a legitimate second-best, provided you label the provenance.",{"title":749,"text":750},"The mobile application's backend","Often a clean JSON API. Check the terms before relying on it — this is where the legal question stops being theoretical, and course six covers it.",{"title":752,"text":753},"Asking","Genuinely underused. Some retailers will hand you a feed if you explain what you need and why. A no costs you an email.",{"type":125,"variant":229,"title":755,"text":756},"Label the substitute, always","If prices for one retailer come from a marketplace listing rather than their own site, that belongs in the data as a provenance field, not in a comment in the code. Six months later nobody will remember, and a slightly different number will be read as a price change rather than a sourcing difference.",{"type":79,"title":758,"paragraphs":759},"How to write a coverage gap so it is useful",[760,761,762,763],"A good gap report is three sentences and a number. What is missing, what was tried, what it would take, and what it costs to leave it.","\"We cannot read this retailer. Nine hundred pages attempted over two days, all refused at the network level before any content was returned; the pattern is categorical rather than rate-related. Getting through would mean a per-page cost roughly four times our current average, with no guarantee of durability. Leaving it means our price position excludes one of eleven competitors in that segment, which matters most on garden furniture where they are the volume leader.\"","That is an input to a decision. Compare it to \"the scraper for this site doesn't work\", which is an apology and tells nobody anything.","The habit underneath all of it: publish what you measured, not what you expected. A page that honestly says the storefront did not return readable product data is worth more than one that pads the gap with a plausible number, because the first can be acted on and the second quietly poisons everything downstream of it.",{"type":84,"title":765,"intro":766,"headers":767,"rows":768},"Worked example: pricing the three answers on one site","A competitor that made up 11% of the feed has gone behind a protection layer. Laid out like this the decision takes an hour. Left undocumented it takes a quarter, and the answer arrived at is usually the one that is not in this table.",[418,719,723,542],[769,774,779,784,789,794],[770,771,772,773],"What it means here","Escalate to rendering plus a residential address","Take the price from the marketplace listing instead","Remove the site and document the gap",[775,776,777,778],"Effort","about two days, then ongoing","half a day","an hour",[780,781,782,783],"Running cost","about 25× the base page rate, on 11% of the feed","base rate, different source","nothing",[785,786,787,788],"What the data becomes","unchanged","a third-party seller's price, not the retailer's — label it","a named absence",[790,791,792,793],"The risk you are taking","it breaks again at the next change, now at 25×","the substitute gets treated as equivalent by everyone downstream","somebody presents a market view with a hole in it",[795,796,797,798],"When it wins","the site sets the market price on your top lines","the substitute is genuinely comparable","the site was a benchmark rather than a threat",{"type":130,"title":168,"intro":800,"items":801},"The three answers are all defensible. What costs money is the fourth thing people do instead.",[802,803,804,805,806,807],"Choosing none of them. The default is a scraper half-fixed every few weeks forever, which costs more than any single column in the table.","Arguing about effort before valuing the data. The question is which decision gets worse without this site, and if the answer is none, the right-hand column is free.","Routing around without labelling the substitute. A marketplace price and a retailer price are different things, and once they share a column nobody can separate them again.","Escalating permanently to solve a temporary problem. The expensive rung tends to stay switched on long after the reason for switching it on has gone.","Writing the coverage gap as an apology. \"It doesn't work\" is not a decision input. \"This competitor is 11% of the feed, unavailable since 14 March, substitute available at marketplace level\" is.","Scheduling no review. A gap documented once and never revisited becomes permanent through neglect rather than through a decision.",{"type":130,"title":809,"intro":810,"items":811},"The maintenance routine worth having","None of this is clever. All of it compounds.",[812,813,814,815,816],"Review the row-count deltas weekly. Ten minutes, and it is the whole early warning system.","Refresh URL lists monthly — content drift is constant and silent.","Hand-check twenty rows against live pages monthly.","Keep a one-line log per source of what broke and what fixed it. It is how a new person becomes useful in a week instead of a quarter.","Re-run the value question annually. Sources that were load-bearing stop being load-bearing and nobody notices until someone audits the bill.",[818,819,820,821],"Three answers: repair, route around, stop. Choosing none of them is the expensive default.","Value the data first — what decision gets worse without it — then cap the effort before you start.","Sitemaps, marketplaces, aggregators, mobile backends and simply asking are all real alternatives. Label the substitute in the data.","A specific coverage gap with numbers is a decision input. \"It doesn't work\" is an apology.",{"text":823,"label":824,"url":825},"Collected data is only useful once rows from different sites can be matched to the same product. That is the next course.","Course 5: matching products across sites","/learn/matching-products-across-sites",[827,880,920,959,967,1005],{"order":828,"slug":829,"title":830,"subtitle":831,"cardText":832,"level":833,"time":834,"lessonCount":835,"lessons":836},1,"competitor-price-monitoring","Build a competitor price monitoring pipeline","Price monitoring looks like a scraping problem for about a week. Then you discover that scraping was the easy part, and the project actually lives or dies on which competitors you picked, whether their listings are really the same product as yours, and whether anyone notices the morning the feed comes back half empty. This course is those eight decisions, in the order you have to make them.","From \"we check three competitors by hand on Mondays\" to a feed you trust enough to reprice from. The eight decisions in order, including the two that quietly ruin most projects.","No coding required","8 lessons, about 90 minutes",8,[837,843,848,854,859,865,870,875],{"slug":838,"navTitle":839,"title":840,"summary":841,"time":842,"needsAccount":72},"what-is-competitor-price-monitoring","What it actually is","What competitor price monitoring actually is","The four stages of a price pipeline, why only two of them are scraping, and the one question to ask before you build anything.","9 min",{"slug":844,"navTitle":845,"title":846,"summary":847,"time":312,"needsAccount":72},"choose-competitors-and-skus","Choosing what to track","Choosing which competitors and which SKUs to track","How to build a list that is small enough to afford and large enough to matter, using margin at risk rather than gut feel.",{"slug":849,"navTitle":850,"title":851,"summary":852,"time":195,"needsAccount":853},"find-competitor-product-urls","Finding product URLs","Finding every competitor product URL without copying them by hand","Four ways to get a competitor's full product URL list, ranked by how much work they are, and what to do when none of them work.",true,{"slug":855,"navTitle":856,"title":857,"summary":858,"time":195,"needsAccount":853},"extract-price-stock-and-shipping","Extracting the fields","Getting price, stock and shipping off the page","Which fields to extract, why the sale price is two fields and not one, and the four ways a price appears on a page.",{"slug":860,"navTitle":861,"title":862,"summary":863,"time":864,"needsAccount":72},"match-listings-to-your-catalogue","Matching to your catalogue","Matching competitor listings to your own catalogue","The stage that decides whether your feed is intelligence or fiction, and the denominator trick that makes bad match rates look good.","13 min",{"slug":866,"navTitle":867,"title":868,"summary":869,"time":312,"needsAccount":72},"schedule-runs-and-catch-silent-failure","Scheduling and data quality","Scheduling runs and catching silent data loss","How often to actually check, and the four alerts that catch a degrading feed before someone reprices from it.",{"slug":871,"navTitle":872,"title":873,"summary":874,"time":842,"needsAccount":853},"export-to-sheets-bi-and-erp","Getting the data out","Getting the data into Sheets, BI or your ERP","Four delivery routes ranked by how likely they are to actually get used, and the column contract that stops downstream jobs breaking.",{"slug":876,"navTitle":877,"title":878,"summary":879,"time":195,"needsAccount":72},"turn-price-data-into-repricing-rules","From data to decisions","Turning price data into repricing decisions","Why \"match the cheapest\" destroys margin, what a rule needs besides a competitor price, and how to start without automating anything.",{"order":881,"slug":882,"title":883,"subtitle":884,"cardText":885,"level":886,"time":887,"lessonCount":888,"lessons":889},2,"ai-agent-web-data-mcp","Give your AI agent live web data via MCP","Ask an assistant what a product costs today and you will usually get a number. It is often wrong, and it is always wrong in the same way: the model is reconstructing a plausible price from training data rather than looking at a page. This course is about closing that gap properly — what the Model Context Protocol actually is, how to wire a server into a client, how to design tools a model can use without hand-holding, and what to put in place before an agent spends your money.","Your agent is confidently wrong about prices because it has never seen one. What MCP is, how to connect a server, how to design tools a model can actually use, and the guardrails you need before you let it loose.","Comfortable editing a config file","6 lessons, about 60 minutes",6,[890,895,900,905,910,915],{"slug":891,"navTitle":892,"title":893,"summary":894,"time":842,"needsAccount":72},"what-is-mcp","What MCP is","What MCP actually is, in plain terms","The Model Context Protocol described without jargon: what problem it solves, its three primitives, and when it is the wrong tool.",{"slug":896,"navTitle":897,"title":898,"summary":899,"time":71,"needsAccount":72},"why-agents-get-live-data-wrong","Why agents get it wrong","Why your agent's answer about a price is wrong","Four distinct failure modes that all look identical from the outside, and how to tell which one you have before you try to fix it.",{"slug":901,"navTitle":902,"title":903,"summary":904,"time":71,"needsAccount":72},"connect-an-mcp-server","Connecting a server","Connecting an MCP server and proving it works","The config for local and remote servers, the four things that go wrong, and how to verify the tools registered rather than assuming.",{"slug":906,"navTitle":907,"title":908,"summary":909,"time":312,"needsAccount":72},"design-tools-an-agent-can-use","Designing usable tools","Designing tools an agent can actually use","A connected server is not a useful server. The model only sees your tool names, descriptions and parameter schemas, so those three things are the entire user interface. Here is what makes a tool get called correctly and what makes it get ignored.",{"slug":911,"navTitle":912,"title":913,"summary":914,"time":195,"needsAccount":853},"give-an-agent-a-scraper","Giving an agent a scraper","Giving an agent a real price feed","A worked example. Connect the ScrapeWise MCP server to a client, let the agent read a live scraper's output, and watch where the hand-off between \"the data is right\" and \"the answer is right\" actually breaks.",{"slug":916,"navTitle":917,"title":918,"summary":919,"time":312,"needsAccount":72},"guardrails-cost-and-untrusted-content","Guardrails and cost","Guardrails, cost control and untrusted content","Live web access turns an agent into something that can spend money and read text written by strangers. Neither is a reason not to do it. Both are reasons to put limits in before you need them.",{"order":921,"slug":922,"title":923,"subtitle":924,"cardText":925,"level":926,"time":927,"lessonCount":888,"lessons":928},3,"product-data-api","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.","Comfortable with HTTP and JSON","6 lessons, about 70 minutes",[929,934,939,944,949,954],{"slug":930,"navTitle":931,"title":932,"summary":933,"time":71,"needsAccount":72},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.",{"slug":935,"navTitle":936,"title":937,"summary":938,"time":71,"needsAccount":72},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":940,"navTitle":941,"title":942,"summary":943,"time":195,"needsAccount":853},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.",{"slug":945,"navTitle":946,"title":947,"summary":948,"time":312,"needsAccount":72},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.",{"slug":950,"navTitle":951,"title":952,"summary":953,"time":312,"needsAccount":72},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"slug":955,"navTitle":956,"title":957,"summary":958,"time":195,"needsAccount":853},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",{"order":5,"slug":4,"title":17,"subtitle":18,"cardText":8,"level":6,"time":7,"lessonCount":888,"lessons":960},[961,962,963,964,965,966],{"slug":67,"navTitle":68,"title":69,"summary":70,"time":71,"needsAccount":72},{"slug":191,"navTitle":192,"title":193,"summary":194,"time":195,"needsAccount":72},{"slug":308,"navTitle":309,"title":310,"summary":311,"time":312,"needsAccount":72},{"slug":447,"navTitle":448,"title":449,"summary":450,"time":312,"needsAccount":72},{"slug":570,"navTitle":571,"title":572,"summary":573,"time":312,"needsAccount":72},{"slug":698,"navTitle":699,"title":700,"summary":701,"time":71,"needsAccount":72},{"order":968,"slug":969,"title":970,"subtitle":971,"cardText":972,"level":973,"time":927,"lessonCount":888,"lessons":974},5,"matching-products-across-sites","Match the same product across different sites","A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.","Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.","You have data from more than one site",[975,980,985,990,995,1000],{"slug":976,"navTitle":977,"title":978,"summary":979,"time":71,"needsAccount":72},"why-matching-is-the-hard-part","Why matching is hard","Why matching is the hard part","The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.",{"slug":981,"navTitle":982,"title":983,"summary":984,"time":312,"needsAccount":72},"identifiers-first-gtin-ean-mpn","Identifiers first","Identifiers first: GTIN, EAN, UPC and MPN","What each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.",{"slug":986,"navTitle":987,"title":988,"summary":989,"time":195,"needsAccount":72},"when-there-is-no-barcode","No barcode","Matching when there is no barcode","Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.",{"slug":991,"navTitle":992,"title":993,"summary":994,"time":195,"needsAccount":72},"variants-bundles-and-multipacks","Variants and packs","Variants, bundles and multipacks","The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.",{"slug":996,"navTitle":997,"title":998,"summary":999,"time":312,"needsAccount":853},"score-confidence-and-build-a-review-queue","Confidence and review","Confidence scores and a review queue worth using","Why one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.",{"slug":1001,"navTitle":1002,"title":1003,"summary":1004,"time":312,"needsAccount":72},"measure-your-match-rate-honestly","Measure it honestly","Measure your match rate honestly","The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.",{"order":888,"slug":1006,"title":1007,"subtitle":1008,"cardText":1009,"level":1010,"time":1011,"lessonCount":968,"lessons":1012},"web-scraping-legal-and-ethical","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.","No legal background assumed","5 lessons, about 55 minutes",[1013,1018,1023,1028,1033],{"slug":1014,"navTitle":1015,"title":1016,"summary":1017,"time":71,"needsAccount":72},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.",{"slug":1019,"navTitle":1020,"title":1021,"summary":1022,"time":312,"needsAccount":72},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.",{"slug":1024,"navTitle":1025,"title":1026,"summary":1027,"time":312,"needsAccount":72},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"slug":1029,"navTitle":1030,"title":1031,"summary":1032,"time":312,"needsAccount":72},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.",{"slug":1034,"navTitle":1035,"title":1036,"summary":1037,"time":195,"needsAccount":72},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.",1791047866703]