[{"data":1,"prerenderedAt":85},["ShallowReactive",2],{"$fy5X7nS4OJp5h3YDZ_Tmhwg2qCGCiRyW9pS_kvCyGUJg":3},{"title":4,"date":5,"dateModified":6,"datePublished":7,"dateModifiedISO":7,"image":8,"content":9,"faq":10,"metaTitle":30,"metaDescription":31,"author":32,"authorBio":33,"authorLinkedin":33,"authorTitle":33,"authorPhoto":33,"lastReviewed":33,"researchBasis":33,"category":34,"readingTime":35,"related":36,"prev":52,"next":33,"toc":55,"takeaways":84},"Build a Web Scraper with Claude and MCP: A Step-by-Step Walkthrough","30 Sep 2026","30 SEP 2026","2026-09-30","/img/news/build-web-scraper-with-claude-mcp-2026.png","\u003Cp>Most of our recent competitor price setups were not built by a person clicking through the portal. They were built by Claude, connected to Scrapewise through our MCP server, with a person answering the questions that needed a decision.\u003C/p>\n\u003Cp>This walkthrough shows how to do the same with your own account. You connect your LLM once, then ask it in plain language to find a site&#39;s price source, build the scraper, test it on one page, run it and read the rows back. The model does the clicking. You keep the decisions and the wallet.\u003C/p>\n\u003Ch2 id=\"what-mcp-changes\">What MCP Changes\u003C/h2>\n\u003Cp>\u003Ca href=\"https://modelcontextprotocol.io/docs/getting-started/intro\">MCP (Model Context Protocol)\u003C/a> is an open standard that lets an AI app call tools on another service. Scrapewise exposes its REST API as MCP tools at one address, \u003Ccode>https://mcp.scrapewise.ai/mcp\u003C/code>. The client discovers the tools by itself, so you do not write any glue code.\u003C/p>\n\u003Cp>Without it, an LLM can only tell you how to build a scraper. With it, the LLM can build one in your account, run it, and check the result.\u003C/p>\n\u003Caside class=\"article__usecase-card\">\u003Cdiv class=\"article__usecase-label\">Related use case\u003C/div>\u003Ch3 class=\"article__usecase-title\">Any-site data scraper\u003C/h3>\u003Cp class=\"article__usecase-blurb\">No-code extraction from any website. Managed infrastructure, no anti-bot headaches.\u003C/p>\u003Ca class=\"article__usecase-link\" href=\"/use-cases/data-scraper\">See how it works →\u003C/a>\u003C/aside>\u003Ch2 id=\"step-1-create-an-api-key\">Step 1: Create an API Key\u003C/h2>\n\u003Cp>In the portal, open \u003Cstrong>Settings &gt; API Keys\u003C/strong> and create a key. Pick the scope:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Ccode>LLM_READ\u003C/code> can list scrapers and groups, preview pages and read your data. It cannot create, run or delete anything. Start with this one if you only want to try it.\u003C/li>\n\u003Cli>\u003Ccode>LLM_FULL\u003C/code> can also create scrapers, change them, run them and delete things. Building scrapers needs this one.\u003C/li>\n\u003C/ul>\n\u003Cp>A normal \u003Ccode>USER\u003C/code> key is refused by the MCP gateway on purpose. MCP is a separate door with its own keys.\u003C/p>\n\u003Ch2 id=\"step-2-connect-your-client\">Step 2: Connect Your Client\u003C/h2>\n\u003Cp>Our docs walk through three Claude clients. For Claude Code it is one command:\u003C/p>\n\u003Cdiv class=\"code-block\">\u003Cbutton type=\"button\" class=\"code-block__copy\" data-copy-code aria-label=\"Copy code\">Copy\u003C/button>\u003Cpre>\u003Ccode class=\"language-bash\">claude mcp add scrapewise --transport http https://mcp.scrapewise.ai/mcp \\\n  --header &#39;Authorization: Bearer &lt;your key&gt;&#39;\n\u003C/code>\u003C/pre>\u003C/div>\n\u003Cp>For Claude Desktop you add the same URL and header to the \u003Ccode>mcpServers\u003C/code> block of its config file. For claude.ai you add a custom connector under \u003Cstrong>Settings &gt; Connectors\u003C/strong> with the URL and an \u003Ccode>Authorization\u003C/code> header. The \u003Ca href=\"https://docs.scrapewise.ai/getting-started/quickstart-mcp\">MCP quickstart\u003C/a> has the exact screens.\u003C/p>\n\u003Cp>Other MCP clients work the same way: the URL plus a \u003Ccode>Bearer\u003C/code> header with your key. ChatGPT and several code editors support MCP too, but we only document the Claude clients, so check that your client can send a custom header before you rely on it.\u003C/p>\n\u003Cp>To check the connection, ask: &quot;List my Scrapewise scraper groups.&quot; If a list comes back, you are connected.\u003C/p>\n\u003Caside class=\"article__inline-cta\">\u003Cp class=\"article__inline-cta-text\">Try ScrapeWise on your own URL. \u003Cstrong>Your first 5 requests are free.\u003C/strong>\u003C/p>\u003Ca class=\"article__inline-cta-btn\" href=\"https://portal.scrapewise.ai/login\" target=\"_blank\" rel=\"noopener\">Start Free →\u003C/a>\u003C/aside>\u003Ch2 id=\"step-3-give-the-model-a-real-brief\">Step 3: Give the Model a Real Brief\u003C/h2>\n\u003Cp>The quality of the result depends on the brief more than on anything else. A good brief for a price scraper names:\u003C/p>\n\u003Cul>\n\u003Cli>The competitor sites and markets.\u003C/li>\n\u003Cli>What one row should be: one product, one size, or one link from your list.\u003C/li>\n\u003Cli>The columns you want, with the same names on every site.\u003C/li>\n\u003Cli>What you do not want: for example no scheduled runs yet, and no deletes.\u003C/li>\n\u003Cli>A spend cap for the session.\u003C/li>\n\u003C/ul>\n\u003Cp>Here is a shortened version of a brief we have used:\u003C/p>\n\u003Cblockquote>\n\u003Cp>Build daily price scrapers for these three shops in the group &quot;Competitor prices&quot;. Look for a JSON source first (the page&#39;s own API calls, Shopify or WooCommerce JSON, JSON-LD) and use HTML selectors only if there is none. Columns on every scraper: name, sku, price, sale_price, currency, pack_grams, url, image, availability. Do not compute anything in the extraction, only copy raw values. Do one test scrape per scraper and show me 5 rows before any full run. Do not schedule anything. Do not delete anything. Stop and ask if a run would cost more than €1.\u003C/p>\n\u003C/blockquote>\n\u003Cp>The line about not computing anything is there for a reason. We asked an extraction model to calculate a price per 50 g once, and it was right on 18 of 69 rows. The story is in \u003Ca href=\"/blogs/llm-data-extraction-rules-do-the-maths-2026\">let AI read the values and let rules do the maths\u003C/a>.\u003C/p>\n\u003Ch2 id=\"step-4-what-the-model-actually-calls\">Step 4: What the Model Actually Calls\u003C/h2>\n\u003Cp>It helps to know the tools, so you can read what the model is doing. These are the ones a scraper build uses most. The full list is in the \u003Ca href=\"https://docs.scrapewise.ai/mcp/tools\">tool catalog\u003C/a>.\u003C/p>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Stage\u003C/th>\n\u003Cth>Tool\u003C/th>\n\u003Cth>What it does\u003C/th>\n\u003C/tr>\n\u003C/thead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Look at a site\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_preview_scraper_from_url\u003C/code>\u003C/td>\n\u003Ctd>Fetches a page and suggests scraper configs with a data preview\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Try an API call\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_preview_scraper_from_curl\u003C/code>\u003C/td>\n\u003Ctd>Same, starting from a curl command the model found in the page&#39;s network calls\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Build\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_create_scraper_group\u003C/code>, \u003Ccode>scrapewise_create_scraper_v2\u003C/code>\u003C/td>\n\u003Ctd>Creates the group and the scraper\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Give it links\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_create_scraper_site\u003C/code>\u003C/td>\n\u003Ctd>Stores the scraper&#39;s link list, with a title per link\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Test once\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_get_scraper_sample_data\u003C/code>\u003C/td>\n\u003Ctd>Runs the scraper on a small first batch and returns rows\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Test a rule\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_preview_scraper_preview_rule\u003C/code>\u003C/td>\n\u003Ctd>Tries one after-scrape rule on a sample value\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Your own prices\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_create_file_scraper\u003C/code>\u003C/td>\n\u003Ctd>Turns your CSV, JSON Lines or Excel file into an upload scraper, free\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Run\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_run_scraper\u003C/code>, \u003Ccode>scrapewise_run_scraper_group\u003C/code>\u003C/td>\n\u003Ctd>Starts a run\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Check\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_get_scraper_load_history\u003C/code>, \u003Ccode>scrapewise_get_scraper_job_errors\u003C/code>\u003C/td>\n\u003Ctd>Run status and the per-link errors\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Read\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_get_scraper_data_group\u003C/code>, \u003Ccode>scrapewise_export_scraper_data_group\u003C/code>\u003C/td>\n\u003Ctd>Pages of rows, or an Excel export\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Schedule\u003C/td>\n\u003Ctd>\u003Ccode>scrapewise_update_scraper_group_schedule\u003C/code>\u003C/td>\n\u003Ctd>Sets the group&#39;s recurring schedule\u003C/td>\n\u003C/tr>\n\u003C/tbody>\u003C/table>\n\u003Cp>For a plain list of addresses there is one shortcut, sized for big lists: \u003Ccode>scrapewise_run_scraper_url_list\u003C/code> creates one automation, stores the whole list and starts one run. The gateway tells every model to use one automation per list, never one per address, and it repeats that in the tool descriptions. Models still try the other way now and then, so it is worth saying it in your brief as well. One call takes up to 5,000 URLs. For more, the REST API takes up to 100,000 URLs per JSON Lines upload, and one scraper can also be fed its links automatically from another scraper&#39;s rows, up to about 700,000 links.\u003C/p>\n\u003Ch2 id=\"step-5-test-on-one-page-then-run\">Step 5: Test on One Page, Then Run\u003C/h2>\n\u003Cp>The sample scrape is the cheapest check there is. It runs the scraper on its first batch and returns rows without writing a run. We ask the model to compare 5 sample rows against the live page by hand: price, before-price, pack size, currency, stock.\u003C/p>\n\u003Cp>Only after that does the model start a full run, and for big sites one scraper at a time. When the run finishes, ask it to check the rows against what it expected: row count against the site&#39;s own total, how many rows have each column filled, and duplicates on the product id. A run that returns 200 OK on every page can still hold nothing useful, as we found when we \u003Ca href=\"/blogs/http-200-not-success-eu-marketplaces-2026\">measured eight European marketplaces\u003C/a>.\u003C/p>\n\u003Ch2 id=\"step-6-after-scrape-rules\">Step 6: After-Scrape Rules\u003C/h2>\n\u003Cp>Once the raw columns are right, the model adds rules on the scraper. The rule kinds today:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Convert currency\u003C/strong> writes an EUR column from a price column, with the ECB rate of the day.\u003C/li>\n\u003Cli>\u003Cstrong>Replace values from a list\u003C/strong> maps each shop&#39;s stock text to one set of values.\u003C/li>\n\u003Cli>\u003Cstrong>Clean up text\u003C/strong> keeps only the characters you choose, for example digits for an EAN.\u003C/li>\n\u003Cli>\u003Cstrong>Take a number from text\u003C/strong> reads the number in front of a unit, such as &quot;100 g&quot;, and can convert kg to g.\u003C/li>\n\u003Cli>\u003Cstrong>Price per unit\u003C/strong> rescales a price to a fixed pack size, for example price per 50 g.\u003C/li>\n\u003C/ul>\n\u003Cp>The model can test each one with the rule preview tool before saving. We used exactly these to build a \u003Ca href=\"/blogs/price-per-unit-comparison-pack-sizes-2026\">price per unit comparison across pack sizes\u003C/a>.\u003C/p>\n\u003Ch2 id=\"guardrails-we-use\">Guardrails We Use\u003C/h2>\n\u003Cp>\u003Cstrong>Start read-only.\u003C/strong> Give the model an \u003Ccode>LLM_READ\u003C/code> key while you learn what it does. Switch to \u003Ccode>LLM_FULL\u003C/code> when you want it to build.\u003C/p>\n\u003Cp>\u003Cstrong>Deletes take two steps.\u003C/strong> Delete tools come with a preview call that returns a token, and the delete only happens with that token within 5 minutes. Claude Desktop and claude.ai also ask before a tool marked destructive runs, by default. We still write &quot;do not delete anything&quot; in the brief.\u003C/p>\n\u003Cp>\u003Cstrong>Scraped text is marked as scraped.\u003C/strong> When a tool returns page content, it arrives wrapped as \u003Ccode>type: &quot;scraped&quot;\u003C/code>, so the client can treat it as untrusted input and not as instructions. A product page that says &quot;ignore your previous instructions&quot; is still just a product page.\u003C/p>\n\u003Cp>\u003Cstrong>The wallet is the limit.\u003C/strong> Scrapewise is pay per page delivered: from €0.15 per 1,000 plain pages, €0.75 with a browser, and up to €3.75 for the hardest sites (see \u003Ca href=\"/pricing\">pricing\u003C/a>). Listing, reading and file uploads are free. When the wallet is empty, a fetching tool returns \u003Ccode>HTTP 402\u003C/code> with \u003Ccode>WALLET_INSUFFICIENT_BALANCE\u003C/code>, and the model should stop and tell you, not retry.\u003C/p>\n\u003Cp>\u003Cstrong>Ask for backups.\u003C/strong> Before the model changes a working scraper, have it read the full config and keep a copy. Our agents do this before every change, and it has saved us more than once.\u003C/p>\n\u003Ch2 id=\"what-went-wrong-when-we-did-this\">What Went Wrong When We Did This\u003C/h2>\n\u003Cp>The model is quick at the build and less careful at the check. Three examples from our own runs.\u003C/p>\n\u003Cp>It reported a setup as complete on more links than an independent check with stricter rules found. One link in the customer&#39;s file had never been built, and two shops showed a before-price equal to today&#39;s price, which is not a sale. Ask for the count against the source file, not against the model&#39;s own list.\u003C/p>\n\u003Cp>It fed every variant URL of one shop into a scraper, and every variant URL returned the same page. Fetch three variant URLs first and compare them.\u003C/p>\n\u003Cp>It trusted a site search as a full catalogue. The search returned 245,000 rows but only about 91,800 different products. More on that in \u003Ca href=\"/blogs/web-scraping-mistakes-price-scrapers-2026\">mistakes we made building price scrapers\u003C/a>.\u003C/p>\n\u003Cp>None of these are MCP problems. They are the same mistakes a person makes, made faster. The fix is the same too: a short checklist the model has to report against after every run.\u003C/p>\n\u003Cp>We are publishing the runbook our agents follow as a skill you can give to your own model: \u003Ca href=\"https://docs.scrapewise.ai/skills/build-scrapers\">build scrapers skill\u003C/a>. If you want to see the output of such a build first, the \u003Ca href=\"/blogs/competitor-price-monitoring-case-study-2026\">competitor price monitoring case study\u003C/a> shows one end to end. Every new account gets 5 free requests, enough to connect and try the preview tools.\u003C/p>\n",{"title":11,"description":12,"badge":13,"benefits":14},"Frequently asked questions","Building scrapers with Claude and MCP: questions answered","FAQ",[15,18,21,24,27],{"title":16,"description":17},"Can Claude build a web scraper for me?","Yes. Connected to the Scrapewise MCP server with a full-access key, Claude can preview a site, create the scraper, add links, run a test scrape, start runs and read the rows back in your account.",{"title":19,"description":20},"Which API key scope do I need?","Use an LLM_READ key to list, preview and read data only. Use an LLM_FULL key when you want the model to create, run or change scrapers. Normal USER keys are refused by the MCP gateway.",{"title":22,"description":23},"Which clients are supported?","Our docs walk through Claude Code, Claude Desktop and claude.ai. Other MCP clients connect the same way, with the gateway URL and an Authorization header that carries your key.",{"title":25,"description":26},"What does it cost to let an AI agent build scrapers?","Listing, reading and file uploads are free. Tools that fetch pages are charged per page delivered, from 0.15 euro per 1,000 plain pages, and every new account gets 5 free requests.",{"title":28,"description":29},"How do I stop the model from deleting things?","Start with a read-only key and say so in your brief. Delete tools also need a separate preview call that returns a token valid for 5 minutes, and Claude Desktop and claude.ai ask before destructive tools run, by default.","Build a Web Scraper with Claude and MCP, Step by Step","Connect Claude to Scrapewise through MCP and let it build, test and run price scrapers for you. The setup, the prompts and the guardrails we use.","Raivo Kartau",null,"Scraping",8,[37,42,47],{"slug":38,"title":39,"image":40,"date":5,"category":34,"excerpt":41},"find-hidden-json-api-shop-page-2026","How to Find the Hidden JSON API Behind a Shop Page (and Why It Is Cheaper)","/img/news/find-hidden-json-api-shop-page-2026.png","Most shop pages load prices from JSON first. How we find those endpoints (Shopify, WooCommerce, Magento, JSON-LD, search APIs) and test them.",{"slug":43,"title":44,"image":45,"date":5,"category":34,"excerpt":46},"llm-data-extraction-rules-do-the-maths-2026","LLM Data Extraction: Let the AI Read, Let Rules Do the Maths","/img/news/llm-data-extraction-rules-do-the-maths-2026.png","Our AI extraction got a price per 50 g right on 18 of 69 rows, and never by calculating. Moving the maths into rules fixed it on every row.",{"slug":48,"title":49,"image":50,"date":5,"category":34,"excerpt":51},"web-scraping-mistakes-price-scrapers-2026","14 Web Scraping Mistakes We Made Building Price Scrapers","/img/news/web-scraping-mistakes-price-scrapers-2026.png","Wrong sources, shifted EANs, Excel eating a brand name, AI doing sums. 14 mistakes from two real price monitoring builds, and the fix for each.",{"slug":53,"title":54},"competitor-price-monitoring-case-study-2026","Competitor Price Monitoring Case Study: 100,000 SKUs, 88% Fewer Requests",[56,60,63,66,69,72,75,78,81],{"level":57,"text":58,"id":59},2,"What MCP Changes","what-mcp-changes",{"level":57,"text":61,"id":62},"Step 1: Create an API Key","step-1-create-an-api-key",{"level":57,"text":64,"id":65},"Step 2: Connect Your Client","step-2-connect-your-client",{"level":57,"text":67,"id":68},"Step 3: Give the Model a Real Brief","step-3-give-the-model-a-real-brief",{"level":57,"text":70,"id":71},"Step 4: What the Model Actually Calls","step-4-what-the-model-actually-calls",{"level":57,"text":73,"id":74},"Step 5: Test on One Page, Then Run","step-5-test-on-one-page-then-run",{"level":57,"text":76,"id":77},"Step 6: After-Scrape Rules","step-6-after-scrape-rules",{"level":57,"text":79,"id":80},"Guardrails We Use","guardrails-we-use",{"level":57,"text":82,"id":83},"What Went Wrong When We Did This","what-went-wrong-when-we-did-this",[],1790769637103]