[{"data":1,"prerenderedAt":118},["ShallowReactive",2],{"$ferYK4inTEfMREOJjFEq2jLvN48-JZgnt21_MsZEwwDE":3},{"title":4,"date":5,"dateModified":6,"datePublished":7,"dateModifiedISO":7,"image":8,"content":9,"faq":10,"metaTitle":4,"metaDescription":30,"author":31,"authorBio":32,"authorLinkedin":32,"authorTitle":32,"authorPhoto":32,"lastReviewed":32,"researchBasis":32,"category":33,"readingTime":34,"related":35,"prev":51,"next":54,"toc":57,"takeaways":117},"14 Web Scraping Mistakes We Made Building Price Scrapers","30 Sep 2026","30 SEP 2026","2026-09-30","/img/news/web-scraping-mistakes-price-scrapers-2026.png","\u003Cp>In September 2026 we built two competitor price setups back to back. One covered full competitor catalogues in two EU markets, matched against a catalogue of about 100,000 SKUs, and cut daily requests from 53,717 to 6,302. The other compared a yarn retailer&#39;s competitor links per 50 g, with the price per unit now right on 100% of rows.\u003C/p>\n\u003Cp>Both work now. Both took longer than they should have, and nearly all of the extra time went into fixing our own mistakes. This is the list, grouped by where they happened, with what each one cost and the check we run now.\u003C/p>\n\u003Cp>None of these are exotic. That is the uncomfortable part.\u003C/p>\n\u003Ch2 id=\"mistakes-in-choosing-the-source\">Mistakes in Choosing the Source\u003C/h2>\n\u003Ch3 id=\"1-fetching-every-variant-url-when-they-all-return-the-same-p\">1. Fetching every variant URL when they all return the same page\u003C/h3>\n\u003Cp>One shop has a URL for every size of every product. We fed them all to the scraper. Each returned the same product page, about 14 times per product, and the run produced no rows from them. That was about €4 of pages for nothing, which is cheap as mistakes go, and a day of confusion, which is not.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> fetch three variant URLs of one product and compare the responses. If they are the same, use one URL per product.\u003C/p>\n\u003Ch3 id=\"2-trusting-a-site-search-as-a-full-catalogue\">2. Trusting a site search as a full catalogue\u003C/h3>\n\u003Cp>A site search looked like a cheaper way to list one shop&#39;s whole range. It returned 245,000 rows, but only about 91,800 different products. Deep search pages repeated the same tiles.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> before switching to a new source, diff its product ids against the old source. Row counts are not enough.\u003C/p>\n\u003Ch3 id=\"3-paying-for-a-browser-the-site-did-not-need\">3. Paying for a browser the site did not need\u003C/h3>\n\u003Cp>One site looked like it needed a rendered page with Super (residential proxy). A plain fetch gave the same rows. On our \u003Ca href=\"/pricing\">pricing\u003C/a> that is €3.75 per 1,000 pages against €0.15.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> run a plain fetch and a rendered fetch on the same pages and compare the rows before choosing the tier.\u003C/p>\n\u003Ch3 id=\"4-using-ai-extraction-where-a-json-source-existed\">4. Using AI extraction where a JSON source existed\u003C/h3>\n\u003Cp>On the yarn project we started with AI extraction on every link. Most of those shops had a structured source: Shopify&#39;s variant JSON, the WooCommerce Store API, or microdata in the page. The structured sources were more accurate and cost about a ninth as much. We now use AI only on the under 10% of links where nothing structured has the pack weight. How we look for these sources is in \u003Ca href=\"/blogs/find-hidden-json-api-shop-page-2026\">finding the hidden JSON API behind a shop page\u003C/a>.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> look for a platform JSON endpoint, JSON-LD and microdata before building an AI or selector scraper.\u003C/p>\n\u003Caside class=\"article__usecase-card\">\u003Cdiv class=\"article__usecase-label\">Related use case\u003C/div>\u003Ch3 class=\"article__usecase-title\">Any-site data scraper\u003C/h3>\u003Cp class=\"article__usecase-blurb\">No-code extraction from any website. Managed infrastructure, no anti-bot headaches.\u003C/p>\u003Ca class=\"article__usecase-link\" href=\"/use-cases/data-scraper\">See how it works →\u003C/a>\u003C/aside>\u003Ch2 id=\"mistakes-in-reading-the-data\">Mistakes in Reading the Data\u003C/h2>\n\u003Ch3 id=\"5-reading-json-arrays-by-position\">5. Reading JSON arrays by position\u003C/h3>\n\u003Cp>An endpoint returned sizes, SKUs and EANs as separate arrays. When one variant had no EAN, every EAN after it moved to the wrong size. Nothing failed. The EANs were just attached to the wrong rows.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> prefer endpoints with one object per variant. Compare 5 rows against the live page, field by field.\u003C/p>\n\u003Ch3 id=\"6-assuming-the-api-price-is-the-page-price\">6. Assuming the API price is the page price\u003C/h3>\n\u003Cp>In one market, the API price and the price on the page differed by a constant factor: the VAT difference between two countries. Some discounts also exist only on the page.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> compare the API price with the page price on about 50 products per market. A constant factor becomes a fixed column. A page-only discount is recorded as a known limit, not hidden.\u003C/p>\n\u003Ch3 id=\"7-asking-ai-to-do-arithmetic\">7. Asking AI to do arithmetic\u003C/h3>\n\u003Cp>We asked the extraction step for a price per 50 g. It was right on 18 of 69 rows, and only where the ball was 50 g, so the price was the answer. It never calculated anything. The story is in \u003Ca href=\"/blogs/llm-data-extraction-rules-do-the-maths-2026\">let the AI read and let rules do the maths\u003C/a>.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> the AI copies printed values only. Every calculation is an after-scrape rule. Recompute a sample by hand after the first run.\u003C/p>\n\u003Ch3 id=\"8-treating-a-crossed-out-price-equal-to-today39s-price-as-a-\">8. Treating a crossed-out price equal to today&#39;s price as a sale\u003C/h3>\n\u003Cp>Two shops show a &quot;before&quot; price that is the same as the price you pay. Our first version called those sales.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> a sale price is filled only when it is lower than the regular price.\u003C/p>\n\u003Ch3 id=\"9-reading-stock-text-as-a-quantity\">9. Reading stock text as a quantity\u003C/h3>\n\u003Cp>An accessory page said &quot;40+ pcs in stock&quot;. The AI read that as the number of pieces in the pack, and the row ended up with two different units.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> each accessory row carries exactly one unit, per 100 g or per piece, and the unit comes from the product, not from the stock line.\u003C/p>\n\u003Ch2 id=\"mistakes-with-the-customer39s-own-data\">Mistakes With the Customer&#39;s Own Data\u003C/h2>\n\u003Ch3 id=\"10-cleaning-the-customer39s-excel-file-by-hand\">10. Cleaning the customer&#39;s Excel file by hand\u003C/h3>\n\u003Cp>A brand in the customer&#39;s catalogue is called &quot;100%&quot;. Excel stores that as the number 1 with a percent format. Our cleaning script read the stored value, so about 540 brand cells per market turned into &quot;1&quot;.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> upload the customer&#39;s raw file as an upload-file scraper, which keeps the displayed text. If a cleaning step is unavoidable, compare displayed and stored values per column. Percentages, dates and codes with leading zeros are the usual victims.\u003C/p>\n\u003Ch3 id=\"11-re-uploading-a-file-already-used-in-matching\">11. Re-uploading a file already used in matching\u003C/h3>\n\u003Cp>Once a file is used in matching, a new upload of it only adds rows and fills empty cells, matched by a key column. We re-uploaded a corrected catalogue expecting a clean replacement.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> get the file right before the first matching run. If it has to change later, plan it as a new upload with its own key column.\u003C/p>\n\u003Ch3 id=\"12-different-column-names-on-different-scrapers\">12. Different column names on different scrapers\u003C/h3>\n\u003Cp>A join read \u003Ccode>sku\u003C/code>. One scraper had written \u003Ccode>variantSku\u003C/code>. The join filled 0 rows and reported no error.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> one fixed set of column names for every scraper in the project, decided before the first build: name, brand, sku, productId, variantId, size, colour, price, url, image, and so on.\u003C/p>\n\u003Caside class=\"article__inline-cta\">\u003Cp class=\"article__inline-cta-text\">Try ScrapeWise on your own URL. \u003Cstrong>Your first 5 requests are free.\u003C/strong>\u003C/p>\u003Ca class=\"article__inline-cta-btn\" href=\"https://portal.scrapewise.ai/login\" target=\"_blank\" rel=\"noopener\">Start Free →\u003C/a>\u003C/aside>\u003Ch2 id=\"mistakes-in-how-we-worked\">Mistakes in How We Worked\u003C/h2>\n\u003Ch3 id=\"13-starting-many-large-runs-at-once\">13. Starting many large runs at once\u003C/h3>\n\u003Cp>We started 18 large first runs at the same time. It was not our best morning. Now first runs go one at a time, and each is checked before the next starts: rows against the site&#39;s total, how many rows have each column filled, duplicates on the product id.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> one large run at a time, with a check after each.\u003C/p>\n\u003Ch3 id=\"14-counting-against-our-own-list-instead-of-the-customer39s\">14. Counting against our own list instead of the customer&#39;s\u003C/h3>\n\u003Cp>After the yarn build, the build notes counted more links as complete than an independent check with stricter rules did. One link in the customer&#39;s file had never been built, and two &quot;sales&quot; were the mistake in number 8.\u003C/p>\n\u003Cp>\u003Cstrong>Check now:\u003C/strong> count completeness against the customer&#39;s original file, not the builder&#39;s list, and have someone other than the builder do the count.\u003C/p>\n\u003Ch2 id=\"what-we-changed-in-our-process\">What We Changed in Our Process\u003C/h2>\n\u003Cp>Most of the rework on the multi-market project came from one ordering mistake. We decided the cheapest daily source and the code columns after the first runs, when they belong before the build. Moving that step forward is most of the time we expect to save next time.\u003C/p>\n\u003Cp>The other change is smaller and it catches more. After every first run, we write one table row per scraper: rows, catalogue gap, fill per column, duplicates, verdict. Reading those rows takes a few minutes, and most of the mistakes above show up in them as an empty column or an odd count.\u003C/p>\n\u003Cp>The full build that these came from is written up in the \u003Ca href=\"/blogs/competitor-price-monitoring-case-study-2026\">competitor price monitoring case study\u003C/a> and the \u003Ca href=\"/blogs/price-per-unit-comparison-pack-sizes-2026\">price per unit comparison\u003C/a>. If you run your own scrapers, the \u003Ca href=\"/blogs/http-200-not-success-eu-marketplaces-2026\">status-code trap\u003C/a> and \u003Ca href=\"/blogs/self-healing-scraper-infrastructure-2026\">self-healing scrapers\u003C/a> cover the failures that come after the build. If you would rather someone else made these mistakes for you, \u003Ca href=\"/use-cases/competitor-price-tracking\">competitor price tracking\u003C/a> is the managed version.\u003C/p>\n",{"title":11,"description":12,"badge":13,"benefits":14},"Frequently asked questions","Web scraping mistakes: questions answered","FAQ",[15,18,21,24,27],{"title":16,"description":17},"What is the most common web scraping mistake?","Trusting output that looks complete. A filled column, a high row count or a 200 status can all hide wrong data, so recompute a sample and compare ids against a second source.",{"title":19,"description":20},"How do I know if variant URLs need separate scraping?","Fetch three variant URLs of one product and compare the responses. If they are identical, scrape one URL per product and take the variants from the page data.",{"title":22,"description":23},"Why do EANs end up on the wrong product?","Usually because arrays of sizes and EANs were read by position and one variant had no EAN, which shifts every later value. Use endpoints with one object per variant.",{"title":25,"description":26},"Should I clean a customer's Excel file before uploading?","Upload the raw file if you can. Excel stores some displayed values differently, such as a brand called 100% saved as the number 1, and a cleaning script can read the stored value.",{"title":28,"description":29},"How many first runs should I start at once?","One large run at a time, with a check after each on row counts, fill rates and duplicates. Starting many at once makes failures hard to trace.","Wrong sources, shifted EANs, Excel eating a brand name, AI doing sums. 14 mistakes from two real price monitoring builds, and the fix for each.","Raivo Kartau",null,"Scraping",6,[36,41,46],{"slug":37,"title":38,"image":39,"date":5,"category":33,"excerpt":40},"build-web-scraper-with-claude-mcp-2026","Build a Web Scraper with Claude and MCP: A Step-by-Step Walkthrough","/img/news/build-web-scraper-with-claude-mcp-2026.png","Connect Claude to Scrapewise through MCP and let it build, test and run price scrapers for you. The setup, the prompts and the guardrails we use.",{"slug":42,"title":43,"image":44,"date":5,"category":33,"excerpt":45},"find-hidden-json-api-shop-page-2026","How to Find the Hidden JSON API Behind a Shop Page (and Why It Is Cheaper)","/img/news/find-hidden-json-api-shop-page-2026.png","Most shop pages load prices from JSON first. How we find those endpoints (Shopify, WooCommerce, Magento, JSON-LD, search APIs) and test them.",{"slug":47,"title":48,"image":49,"date":5,"category":33,"excerpt":50},"llm-data-extraction-rules-do-the-maths-2026","LLM Data Extraction: Let the AI Read, Let Rules Do the Maths","/img/news/llm-data-extraction-rules-do-the-maths-2026.png","Our AI extraction got a price per 50 g right on 18 of 69 rows, and never by calculating. Moving the maths into rules fixed it on every row.",{"slug":52,"title":53},"amazon-buy-box-competition-uk-2026","Amazon Buy Box Competition: 72% of Listings With Offer Data Have One Seller",{"slug":55,"title":56},"price-per-unit-comparison-pack-sizes-2026","Price per Unit: How to Compare Competitor Prices Across Pack Sizes",[58,62,66,69,72,75,78,81,84,87,90,93,96,99,102,105,108,111,114],{"level":59,"text":60,"id":61},2,"Mistakes in Choosing the Source","mistakes-in-choosing-the-source",{"level":63,"text":64,"id":65},3,"1. Fetching every variant URL when they all return the same page","1-fetching-every-variant-url-when-they-all-return-the-same-p",{"level":63,"text":67,"id":68},"2. Trusting a site search as a full catalogue","2-trusting-a-site-search-as-a-full-catalogue",{"level":63,"text":70,"id":71},"3. Paying for a browser the site did not need","3-paying-for-a-browser-the-site-did-not-need",{"level":63,"text":73,"id":74},"4. Using AI extraction where a JSON source existed","4-using-ai-extraction-where-a-json-source-existed",{"level":59,"text":76,"id":77},"Mistakes in Reading the Data","mistakes-in-reading-the-data",{"level":63,"text":79,"id":80},"5. Reading JSON arrays by position","5-reading-json-arrays-by-position",{"level":63,"text":82,"id":83},"6. Assuming the API price is the page price","6-assuming-the-api-price-is-the-page-price",{"level":63,"text":85,"id":86},"7. Asking AI to do arithmetic","7-asking-ai-to-do-arithmetic",{"level":63,"text":88,"id":89},"8. Treating a crossed-out price equal to today&#39;s price as a sale","8-treating-a-crossed-out-price-equal-to-today39s-price-as-a-",{"level":63,"text":91,"id":92},"9. Reading stock text as a quantity","9-reading-stock-text-as-a-quantity",{"level":59,"text":94,"id":95},"Mistakes With the Customer&#39;s Own Data","mistakes-with-the-customer39s-own-data",{"level":63,"text":97,"id":98},"10. Cleaning the customer&#39;s Excel file by hand","10-cleaning-the-customer39s-excel-file-by-hand",{"level":63,"text":100,"id":101},"11. Re-uploading a file already used in matching","11-re-uploading-a-file-already-used-in-matching",{"level":63,"text":103,"id":104},"12. Different column names on different scrapers","12-different-column-names-on-different-scrapers",{"level":59,"text":106,"id":107},"Mistakes in How We Worked","mistakes-in-how-we-worked",{"level":63,"text":109,"id":110},"13. Starting many large runs at once","13-starting-many-large-runs-at-once",{"level":63,"text":112,"id":113},"14. Counting against our own list instead of the customer&#39;s","14-counting-against-our-own-list-instead-of-the-customer39s",{"level":59,"text":115,"id":116},"What We Changed in Our Process","what-we-changed-in-our-process",[],1790769637107]