Learn

Six courses on getting web data that is actually usable

Most guides stop at "here is how to scrape a page". The hard part is everything after: picking the right competitors, matching their listings to your own products, noticing the day a run quietly returns half the rows, keeping it alive through a redesign, and being able to answer your legal team when they ask. These courses cover that part. Written, free, no email wall, no video you have to sit through.

  • 6courses
  • 37written lessons
  • €0and no sign-up

What makes these different

Three rules we held ourselves to while writing them.

Nothing is gated

No form, no drip sequence, no "download the PDF". Every lesson is a public page you can read, link to and quote. If a lesson is useful you should be able to send it to a colleague without making them register for anything.

The product is optional

Seven of the thirty-seven lessons walk through doing the thing in Scrapewise, and they say so at the top. The other thirty are method: they work whether you build it yourself, buy it from us, or buy it from somebody else. We would rather you understand the problem than be sold to.

Written by people who got it wrong first

The failure modes in here are ones we hit running production scrapers: a match rate that looked fine until we checked the denominator, a redesign that silently halved a feed, a price read correctly from the wrong country, an agent that confidently quoted a number from a cached page. The warnings are specific because they cost us something.

The courses

Six courses, thirty-seven lessons, roughly seven hours of reading if you did all of it in one go — which nobody does. Each course is standalone and every lesson links to the ones either side of it, so landing on one from search and leaving again is a perfectly good way to use this.

Course 1

Build a competitor price monitoring pipeline

From "we check three competitors by hand on Mondays" to a feed you trust enough to reprice from. The eight decisions in order, including the two that quietly ruin most projects.

  • No coding required
  • 8 lessons, about 90 minutes
Open the course
  1. 01What it actually isThe four stages of a price pipeline, why only two of them are scraping, and the one question to ask before you build anything.9 min
  2. 02Choosing what to trackHow to build a list that is small enough to afford and large enough to matter, using margin at risk rather than gut feel.11 min
  3. 03Finding product URLsFour ways to get a competitor's full product URL list, ranked by how much work they are, and what to do when none of them work.account12 min
  4. 04Extracting the fieldsWhich fields to extract, why the sale price is two fields and not one, and the four ways a price appears on a page.account12 min
  5. 05Matching to your catalogueThe stage that decides whether your feed is intelligence or fiction, and the denominator trick that makes bad match rates look good.13 min
  6. 06Scheduling and data qualityHow often to actually check, and the four alerts that catch a degrading feed before someone reprices from it.11 min
  7. 07Getting the data outFour delivery routes ranked by how likely they are to actually get used, and the column contract that stops downstream jobs breaking.account9 min
  8. 08From data to decisionsWhy "match the cheapest" destroys margin, what a rule needs besides a competitor price, and how to start without automating anything.12 min
Course 2

Give your AI agent live web data via MCP

Your agent is confidently wrong about prices because it has never seen one. What MCP is, how to connect a server, how to design tools a model can actually use, and the guardrails you need before you let it loose.

  • Comfortable editing a config file
  • 6 lessons, about 60 minutes
Open the course
  1. 01What MCP isThe Model Context Protocol described without jargon: what problem it solves, its three primitives, and when it is the wrong tool.9 min
  2. 02Why agents get it wrongFour distinct failure modes that all look identical from the outside, and how to tell which one you have before you try to fix it.10 min
  3. 03Connecting a serverThe config for local and remote servers, the four things that go wrong, and how to verify the tools registered rather than assuming.10 min
  4. 04Designing usable toolsA connected server is not a useful server. The model only sees your tool names, descriptions and parameter schemas, so those three things are the entire user interface. Here is what makes a tool get called correctly and what makes it get ignored.11 min
  5. 05Giving an agent a scraperA worked example. Connect the ScrapeWise MCP server to a client, let the agent read a live scraper's output, and watch where the hand-off between "the data is right" and "the answer is right" actually breaks.account12 min
  6. 06Guardrails and costLive web access turns an agent into something that can spend money and read text written by strangers. Neither is a reason not to do it. Both are reasons to put limits in before you need them.11 min
Course 3

Pull product data over an API

Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.

  • Comfortable with HTTP and JSON
  • 6 lessons, about 70 minutes
Open the course
  1. 01API, scraper or datasetThree ways to get product data, the honest cost of each, and the specific question that decides between them.10 min
  2. 02Auth and the first callBearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.10 min
  3. 03Declaring the fieldsAn extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.account12 min
  4. 04Async runs and pollingWhy collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.11 min
  5. 05Errors and retriesWhich failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.11 min
  6. 06Into your stackScheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.account12 min
Course 4

Keep scrapers alive after the first week

Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.

  • You already have something running
  • 6 lessons, about 65 minutes
Open the course
  1. 01Why scrapers breakBreakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.10 min
  2. 02Read the failureA diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.12 min
  3. 03Durable selectorsA ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.11 min
  4. 04Bot wallsWhat a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.11 min
  5. 05Monitor the feedSix checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.11 min
  6. 06When a site winsA decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.10 min
Course 5

Match the same product across different sites

Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.

  • You have data from more than one site
  • 6 lessons, about 70 minutes
Open the course
  1. 01Why matching is hardThe same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.10 min
  2. 02Identifiers firstWhat each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.11 min
  3. 03No barcodeNormalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.12 min
  4. 04Variants and packsThe highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.12 min
  5. 05Confidence and reviewWhy one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.account11 min
  6. 06Measure it honestlyThe denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.11 min
Course 6

The legal and ethical side, without the hand-waving

The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.

  • No legal background assumed
  • 5 lessons, about 55 minutes
Open the course
  1. 01Is it legal?"Is scraping legal" bundles access, copying and use into a single question. Separating them is most of the work.10 min
  2. 02Terms and loginsWhy a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.11 min
  3. 03Personal dataPublic does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.11 min
  4. 04Conduct and rate limitsThe conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.11 min
  5. 05Briefing legalA one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.12 min

Where to start

You own pricing

  • You are checking competitor prices by hand, or in a spreadsheet somebody updates on Mondays
  • You have a feed already but you do not trust it, and you cannot say why
  • You are about to buy a price monitoring tool and want to know which questions actually separate them
Start with course one →

You are building the pipeline

  • You went looking for a retailer's API documentation and found an application form
  • You want to know what an authenticated run looks like before committing to one
  • You are wiring a third-party feed into an existing stack and want to know where it will break
Start with course three →

Something already runs, and it keeps breaking

  • A redesign took out half your sources and you found out from a colleague
  • You cannot tell a block from an empty result without guessing
  • You suspect a feed is wrong but every job is green
Start with course four →

You have to sign it off

  • Somebody asked "are we allowed to do this?" and the room went quiet
  • You need to brief legal, procurement or a customer's security review
  • You want to know where a vendor's line is before you depend on them
Start with course six →
FAQ

About the courses

The practical questions, answered before you start.

Not for thirty of the thirty-seven lessons. Seven of them walk through building the thing in Scrapewise and are labelled "needs an account" at the top of the page, so you can see what you are getting into before you click. A new account starts with five free requests and no card, which is enough to follow those lessons.

Rather have us set it up?

The courses cover doing it yourself. If you would rather hand over the list of competitors and get the feed back, that is the other half of what we do.