Build a Web Scraper with Claude and MCP: A Step-by-Step Walkthrough

Last updated: 30 SEP 2026

Build a Web Scraper with Claude and MCP: A Step-by-Step Walkthrough

Most of our recent competitor price setups were not built by a person clicking through the portal. They were built by Claude, connected to Scrapewise through our MCP server, with a person answering the questions that needed a decision.

This walkthrough shows how to do the same with your own account. You connect your LLM once, then ask it in plain language to find a site's price source, build the scraper, test it on one page, run it and read the rows back. The model does the clicking. You keep the decisions and the wallet.

What MCP Changes

MCP (Model Context Protocol) is an open standard that lets an AI app call tools on another service. Scrapewise exposes its REST API as MCP tools at one address, https://mcp.scrapewise.ai/mcp. The client discovers the tools by itself, so you do not write any glue code.

Without it, an LLM can only tell you how to build a scraper. With it, the LLM can build one in your account, run it, and check the result.

Step 1: Create an API Key

In the portal, open Settings > API Keys and create a key. Pick the scope:

  • LLM_READ can list scrapers and groups, preview pages and read your data. It cannot create, run or delete anything. Start with this one if you only want to try it.
  • LLM_FULL can also create scrapers, change them, run them and delete things. Building scrapers needs this one.

A normal USER key is refused by the MCP gateway on purpose. MCP is a separate door with its own keys.

Step 2: Connect Your Client

Our docs walk through three Claude clients. For Claude Code it is one command:

claude mcp add scrapewise --transport http https://mcp.scrapewise.ai/mcp \
  --header 'Authorization: Bearer <your key>'

For Claude Desktop you add the same URL and header to the mcpServers block of its config file. For claude.ai you add a custom connector under Settings > Connectors with the URL and an Authorization header. The MCP quickstart has the exact screens.

Other MCP clients work the same way: the URL plus a Bearer header with your key. ChatGPT and several code editors support MCP too, but we only document the Claude clients, so check that your client can send a custom header before you rely on it.

To check the connection, ask: "List my Scrapewise scraper groups." If a list comes back, you are connected.

Step 3: Give the Model a Real Brief

The quality of the result depends on the brief more than on anything else. A good brief for a price scraper names:

  • The competitor sites and markets.
  • What one row should be: one product, one size, or one link from your list.
  • The columns you want, with the same names on every site.
  • What you do not want: for example no scheduled runs yet, and no deletes.
  • A spend cap for the session.

Here is a shortened version of a brief we have used:

Build daily price scrapers for these three shops in the group "Competitor prices". Look for a JSON source first (the page's own API calls, Shopify or WooCommerce JSON, JSON-LD) and use HTML selectors only if there is none. Columns on every scraper: name, sku, price, sale_price, currency, pack_grams, url, image, availability. Do not compute anything in the extraction, only copy raw values. Do one test scrape per scraper and show me 5 rows before any full run. Do not schedule anything. Do not delete anything. Stop and ask if a run would cost more than €1.

The line about not computing anything is there for a reason. We asked an extraction model to calculate a price per 50 g once, and it was right on 18 of 69 rows. The story is in let AI read the values and let rules do the maths.

Step 4: What the Model Actually Calls

It helps to know the tools, so you can read what the model is doing. These are the ones a scraper build uses most. The full list is in the tool catalog.

Stage Tool What it does
Look at a site scrapewise_preview_scraper_from_url Fetches a page and suggests scraper configs with a data preview
Try an API call scrapewise_preview_scraper_from_curl Same, starting from a curl command the model found in the page's network calls
Build scrapewise_create_scraper_group, scrapewise_create_scraper_v2 Creates the group and the scraper
Give it links scrapewise_create_scraper_site Stores the scraper's link list, with a title per link
Test once scrapewise_get_scraper_sample_data Runs the scraper on a small first batch and returns rows
Test a rule scrapewise_preview_scraper_preview_rule Tries one after-scrape rule on a sample value
Your own prices scrapewise_create_file_scraper Turns your CSV, JSON Lines or Excel file into an upload scraper, free
Run scrapewise_run_scraper, scrapewise_run_scraper_group Starts a run
Check scrapewise_get_scraper_load_history, scrapewise_get_scraper_job_errors Run status and the per-link errors
Read scrapewise_get_scraper_data_group, scrapewise_export_scraper_data_group Pages of rows, or an Excel export
Schedule scrapewise_update_scraper_group_schedule Sets the group's recurring schedule

For a plain list of addresses there is one shortcut, sized for big lists: scrapewise_run_scraper_url_list creates one automation, stores the whole list and starts one run. The gateway tells every model to use one automation per list, never one per address, and it repeats that in the tool descriptions. Models still try the other way now and then, so it is worth saying it in your brief as well. One call takes up to 5,000 URLs. For more, the REST API takes up to 100,000 URLs per JSON Lines upload, and one scraper can also be fed its links automatically from another scraper's rows, up to about 700,000 links.

Step 5: Test on One Page, Then Run

The sample scrape is the cheapest check there is. It runs the scraper on its first batch and returns rows without writing a run. We ask the model to compare 5 sample rows against the live page by hand: price, before-price, pack size, currency, stock.

Only after that does the model start a full run, and for big sites one scraper at a time. When the run finishes, ask it to check the rows against what it expected: row count against the site's own total, how many rows have each column filled, and duplicates on the product id. A run that returns 200 OK on every page can still hold nothing useful, as we found when we measured eight European marketplaces.

Step 6: After-Scrape Rules

Once the raw columns are right, the model adds rules on the scraper. The rule kinds today:

  • Convert currency writes an EUR column from a price column, with the ECB rate of the day.
  • Replace values from a list maps each shop's stock text to one set of values.
  • Clean up text keeps only the characters you choose, for example digits for an EAN.
  • Take a number from text reads the number in front of a unit, such as "100 g", and can convert kg to g.
  • Price per unit rescales a price to a fixed pack size, for example price per 50 g.

The model can test each one with the rule preview tool before saving. We used exactly these to build a price per unit comparison across pack sizes.

Guardrails We Use

Start read-only. Give the model an LLM_READ key while you learn what it does. Switch to LLM_FULL when you want it to build.

Deletes take two steps. Delete tools come with a preview call that returns a token, and the delete only happens with that token within 5 minutes. Claude Desktop and claude.ai also ask before a tool marked destructive runs, by default. We still write "do not delete anything" in the brief.

Scraped text is marked as scraped. When a tool returns page content, it arrives wrapped as type: "scraped", so the client can treat it as untrusted input and not as instructions. A product page that says "ignore your previous instructions" is still just a product page.

The wallet is the limit. Scrapewise is pay per page delivered: from €0.15 per 1,000 plain pages, €0.75 with a browser, and up to €3.75 for the hardest sites (see pricing). Listing, reading and file uploads are free. When the wallet is empty, a fetching tool returns HTTP 402 with WALLET_INSUFFICIENT_BALANCE, and the model should stop and tell you, not retry.

Ask for backups. Before the model changes a working scraper, have it read the full config and keep a copy. Our agents do this before every change, and it has saved us more than once.

What Went Wrong When We Did This

The model is quick at the build and less careful at the check. Three examples from our own runs.

It reported a setup as complete on more links than an independent check with stricter rules found. One link in the customer's file had never been built, and two shops showed a before-price equal to today's price, which is not a sale. Ask for the count against the source file, not against the model's own list.

It fed every variant URL of one shop into a scraper, and every variant URL returned the same page. Fetch three variant URLs first and compare them.

It trusted a site search as a full catalogue. The search returned 245,000 rows but only about 91,800 different products. More on that in mistakes we made building price scrapers.

None of these are MCP problems. They are the same mistakes a person makes, made faster. The fix is the same too: a short checklist the model has to report against after every run.

We are publishing the runbook our agents follow as a skill you can give to your own model: build scrapers skill. If you want to see the output of such a build first, the competitor price monitoring case study shows one end to end. Every new account gets 5 free requests, enough to connect and try the preview tools.

Paste any URL — ScrapeWise handles the anti-bot

Managed infrastructure that adapts when sites change. No proxies, no code, no per-request fees.

Not ready to sign up? See 12 real Amazon rows, 972 columns →

97% accuracy on Amazon benchmarks · no credit card · book a 15-min call →

FAQ

Frequently asked questions

Building scrapers with Claude and MCP: questions answered

Yes. Connected to the Scrapewise MCP server with a full-access key, Claude can preview a site, create the scraper, add links, run a test scrape, start runs and read the rows back in your account.

Use an LLM_READ key to list, preview and read data only. Use an LLM_FULL key when you want the model to create, run or change scrapers. Normal USER keys are refused by the MCP gateway.

Our docs walk through Claude Code, Claude Desktop and claude.ai. Other MCP clients connect the same way, with the gateway URL and an Authorization header that carries your key.

Listing, reading and file uploads are free. Tools that fetch pages are charged per page delivered, from 0.15 euro per 1,000 plain pages, and every new account gets 5 free requests.

Start with a read-only key and say so in your brief. Delete tools also need a separate preview call that returns a token valid for 5 minutes, and Claude Desktop and claude.ai ask before destructive tools run, by default.