Keep scrapers alive after the first week
Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.
What you will be able to do
- Name the five reasons a scraper stops working, and tell them apart from the evidence rather than from a hunch
- Read a failure down to its cause instead of restarting the job and hoping
- Write selectors that survive a front-end rewrite, and recognise the ones that will not
- Understand what a bot wall is actually measuring, and what genuinely changes the outcome
- Monitor the output of a feed, not just whether the job exited zero
- Decide, with a rule rather than a mood, when to stop fighting a site
The six lessons
Read in order the first time. After that it works as a reference — lesson two is the one people come back to.
- 01The five reasons a scraper stops workingBreakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.10 min read Read lesson 1 →
- 02Read the failure, not the symptomA diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.12 min read Read lesson 2 →
- 03Selectors that survive a redesignA ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.11 min read Read lesson 3 →
- 04Bot walls, and what actually changes the outcomeWhat a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.11 min read Read lesson 4 →
- 05Monitor the feed, not the runSix checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.11 min read Read lesson 5 →
- 06Deciding what to do when a site winsA decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.10 min read Read lesson 6 →
Who this is for
Written for
- Anyone who already has scrapers running and keeps getting surprised by them
- Developers who inherited somebody else's collection layer and want to stop firefighting it
- Analysts who depend on a scraped feed and need to know how much to trust it
Not written for
- First-time scraper authors — start with course one or course three, this assumes something already runs
- Anyone looking for techniques to defeat a specific site's protection. That is not what this is, and lesson four explains why that framing loses.
Before you start
What people usually want to know when a scraper has just broken.
Lesson two. It is the diagnosis lesson and it is deliberately the longest. Most wasted maintenance time comes from fixing the wrong thing — rewriting a selector when the page never loaded, or rotating proxies when the product was simply discontinued.
Start at lesson one
Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.