[{"data":1,"prerenderedAt":89},["ShallowReactive",2],{"$fLIGdSN6aXUVPODxpZoVPFMV79MxzwwMuNrT2ZOSq8_s":3},{"title":4,"date":5,"dateModified":6,"datePublished":7,"dateModifiedISO":7,"image":8,"content":9,"faq":10,"metaTitle":30,"metaDescription":31,"author":32,"authorBio":33,"authorLinkedin":33,"authorTitle":33,"authorPhoto":33,"lastReviewed":33,"researchBasis":33,"category":34,"readingTime":35,"related":36,"prev":53,"next":56,"toc":59,"takeaways":88},"Competitor Price Monitoring Case Study: 100,000 SKUs, 88% Fewer Requests","30 Sep 2026","30 SEP 2026","2026-09-30","/img/news/competitor-price-monitoring-case-study-2026.png","\u003Cp>In late September 2026 a retailer asked us for daily competitor prices. The brief was short: the full catalogue of every competitor on their list, two EU markets, one catalogue of about 24,000 models (around 100,000 sizes and colours), prices in local currency and in EUR, and every competitor product matched to their own.\u003C/p>\n\u003Cp>This competitor price monitoring case study is what we built, in the order we built it, including the parts we had to redo. The customer stays anonymous, and so do their competitors. The numbers are from our own logs.\u003C/p>\n\u003Cp>One sentence from our side decided most of the choices. We wanted the smallest daily scraping setup that still gives the most matching information. Every design question below was answered against that.\u003C/p>\n\u003Ch2 id=\"the-setup-in-one-picture\">The Setup in One Picture\u003C/h2>\n\u003Cp>We ended up with five parts. Each one is a normal Scrapewise feature, and no custom code was written for this customer.\u003C/p>\n\u003Col>\n\u003Cli>\u003Cstrong>The master.\u003C/strong> The customer&#39;s own catalogue file, uploaded as an upload-file scraper in each market&#39;s group, so it sits next to the competitor data.\u003C/li>\n\u003Cli>\u003Cstrong>One price scraper per site, run daily.\u003C/strong> The cheapest source that gives a price and a stable product key for the whole catalogue.\u003C/li>\n\u003Cli>\u003Cstrong>One enrichment run per site, run once.\u003C/strong> Item-level data that does not change: EAN, maker codes, size, colour, brand and image, joined to the daily rows on the product key.\u003C/li>\n\u003Cli>\u003Cstrong>One matching config per site.\u003C/strong> EAN first, then any exact code, then brand and model, then size and colour, and text last.\u003C/li>\n\u003Cli>\u003Cstrong>A check after every first run.\u003C/strong> Rows against the site&#39;s own total, fill rates per column, duplicates.\u003C/li>\n\u003C/ol>\n\u003Cp>The split between steps 2 and 3 is where the savings come from. A product&#39;s EAN does not change every day. Its price might. So we only fetch the price every day.\u003C/p>\n\u003Caside class=\"article__usecase-card\">\u003Cdiv class=\"article__usecase-label\">Related use case\u003C/div>\u003Ch3 class=\"article__usecase-title\">Competitor price tracking\u003C/h3>\u003Cp class=\"article__usecase-blurb\">Automated price monitoring across marketplaces. 97% accuracy, no per-SKU fees.\u003C/p>\u003Ca class=\"article__usecase-link\" href=\"/use-cases/competitor-price-tracking\">See how it works →\u003C/a>\u003C/aside>\u003Ch2 id=\"api-first-html-last\">API First, HTML Last\u003C/h2>\n\u003Cp>For each site we opened a category page with the browser&#39;s network tab and looked at the requests the page itself makes. On most modern shops the prices arrive as JSON from a search or listing service, and the HTML is drawn from that JSON afterwards.\u003C/p>\n\u003Cp>What we found, by the kind of source each site offered:\u003C/p>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>What the site offered\u003C/th>\n\u003Cth>Daily price source we chose\u003C/th>\n\u003Cth>One-time enrichment\u003C/th>\n\u003C/tr>\n\u003C/thead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>A search API with price and variant id\u003C/td>\n\u003Ctd>That API, at its largest page size\u003C/td>\n\u003Ctd>Product pages once, for EAN\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>A price API that takes a batch of ids\u003C/td>\n\u003Ctd>Batches of ids, at the largest batch the API accepts\u003C/td>\n\u003Ctd>A details API once\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>HTML listing cheaper than any API\u003C/td>\n\u003Ctd>HTML listing, maximum items per page\u003C/td>\n\u003Ctd>The site&#39;s variant data once\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Only an HTML table\u003C/td>\n\u003Ctd>HTML table listing\u003C/td>\n\u003Ctd>None needed\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>One &quot;from&quot; price per model\u003C/td>\n\u003Ctd>Listing only\u003C/td>\n\u003Ctd>None\u003C/td>\n\u003C/tr>\n\u003C/tbody>\u003C/table>\n\u003Cp>One shop on the list was dropped. It set its market only through a cookie created by a form post, and we could not replay that reliably.\u003C/p>\n\u003Cp>If you have not looked for these endpoints before, we wrote up the method in \u003Ca href=\"/blogs/find-hidden-json-api-shop-page-2026\">how to find the hidden JSON API behind a shop page\u003C/a>. The short version: JSON from the shop&#39;s own services survives a redesign much better than a CSS selector does.\u003C/p>\n\u003Ch2 id=\"how-the-daily-requests-fell-by-88\">How the Daily Requests Fell by 88%\u003C/h2>\n\u003Cp>The first working version was not cheap. It did the job, but it asked each site the same questions more than once. Here is what we changed, one fix at a time.\u003C/p>\n\u003Cp>\u003Cstrong>Largest page and batch sizes.\u003C/strong> One site&#39;s price API took 100 ids per call in our first config. We tried 500 and 1,000. At 1,000 it returned an error, at 500 it worked. That alone cut that site from 2,072 requests to 415 per market.\u003C/p>\n\u003Cp>\u003Cstrong>Removing double fetches.\u003C/strong> Two of our scrapers on one site read the same URLs for different fields. Another site had a list scraper that re-read pages a second scraper already covered. We merged or filtered them.\u003C/p>\n\u003Cp>\u003Cstrong>Plain fetch instead of a browser.\u003C/strong> One site looked like it needed a rendered page with Super (residential proxy). A plain fetch returned the same rows. On \u003Ca href=\"/pricing\">our pricing\u003C/a> that is €0.15 per 1,000 pages instead of €3.75, so a 25 times cheaper page for the same data. We compared a full plain run against a rendered run before switching.\u003C/p>\n\u003Cp>\u003Cstrong>Per-size price or per-model price.\u003C/strong> If every size of a jacket costs the same, you only need one price per model. We tested it instead of assuming it. We took models with two or more sizes and counted how often all sizes had the same price. Depending on the site it was 80% to 94%, never 95%. On one site 7 to 8% of sizes carried a surcharge. So we kept per-size prices where they cost no extra request, and on the site with surcharges we derive the size price from the base price plus the surcharge.\u003C/p>\n\u003Cp>\u003Cstrong>Details that never change become one-time sources.\u003C/strong> EANs, maker codes and images moved out of the daily run completely. That one-time work is where the big page counts are: about 23,000 product pages on one site, and an estimated 190,000 to 250,000 product pages on another, read once instead of every day.\u003C/p>\n\u003Cp>The result, measured on the sites where we changed the daily source: from 53,717 requests a day to 6,302, which is 88% fewer. Our estimate for the whole setup went from about €8.80 a day to about €1.90 a day, plus 10 to 20% for retries and blocked pages.\u003C/p>\n\u003Caside class=\"article__inline-cta\">\u003Cp class=\"article__inline-cta-text\">Try ScrapeWise on your own URL. \u003Cstrong>Your first 5 requests are free.\u003C/strong>\u003C/p>\u003Ca class=\"article__inline-cta-btn\" href=\"https://portal.scrapewise.ai/login\" target=\"_blank\" rel=\"noopener\">Start Free →\u003C/a>\u003C/aside>\u003Ch2 id=\"the-mistakes-that-cost-us-most\">The Mistakes That Cost Us Most\u003C/h2>\n\u003Cp>We lost most of the four build days to rework. Three mistakes did most of the damage.\u003C/p>\n\u003Cp>\u003Cstrong>Per-variant URLs that return the same page.\u003C/strong> One shop has a separate URL for every size. We fed all of them. They all returned the same product page, about 14 times over, and the scraper wrote no rows from them. The fix was one URL per product. Now we always fetch three variant URLs of one product first and compare them.\u003C/p>\n\u003Cp>\u003Cstrong>Trusting a site search as a full catalogue.\u003C/strong> On one shop the site search looked like a cheaper way to list everything. It returned 245,000 rows, but only about 91,800 different products, because the deep pages kept repeating the same tiles. The old category listing stayed. Rule since then: diff the product ids of the new source against the old one before switching.\u003C/p>\n\u003Cp>\u003Cstrong>Assuming the API price is the page price.\u003C/strong> In one market the API price and the shown price differed by a constant factor, the VAT difference between two countries. Once we knew it was constant, a fixed divisor column fixed it. Some discounts only exist on the page and never reach any API. We record those as a known limit instead of pretending they are covered.\u003C/p>\n\u003Cp>There is a longer list, with the smaller ones too, in \u003Ca href=\"/blogs/web-scraping-mistakes-price-scrapers-2026\">mistakes we made building price scrapers\u003C/a>.\u003C/p>\n\u003Ch2 id=\"same-column-names-on-every-scraper\">Same Column Names on Every Scraper\u003C/h2>\n\u003Cp>This one is boring and it matters. Every scraper in the setup writes the same column names: name, brand, sku, productId, variantId, size, colour, category, price, url, image, the raw ean and mpn, and fixed columns for competitor, market and currency.\u003C/p>\n\u003Cp>Then a few after-scrape rules add the working columns. &quot;Convert currency&quot; writes \u003Ccode>priceEur\u003C/code> from \u003Ccode>price\u003C/code> with the ECB rate of the day. &quot;Clean up text&quot; keeps only digits in \u003Ccode>eanClean\u003C/code> and only letters and digits in \u003Ccode>mpnClean\u003C/code>. &quot;Replace values from a list&quot; turns each shop&#39;s stock wording into IN_STOCK, OUT_OF_STOCK, BACKORDER or UNKNOWN.\u003C/p>\n\u003Cp>We learned why this matters the hard way. A join that read \u003Ccode>sku\u003C/code> filled 0 rows, because one scraper had written \u003Ccode>variantSku\u003C/code>. Nothing crashed. The column was just empty.\u003C/p>\n\u003Ch2 id=\"matching-codes-first-text-last\">Matching: Codes First, Text Last\u003C/h2>\n\u003Cp>The daily prices are only useful once each competitor row is tied to the right product in the customer&#39;s catalogue. We kept one order on every site: EAN, then any exact maker code after cleaning, then brand and model, then size and colour, and product names last.\u003C/p>\n\u003Cp>The biggest single win was a code hiding in another column. On one site the shop&#39;s own SKU, with the brand prefix removed, was the maker&#39;s part number. It agreed with the EAN on 99.1% of the rows where both existed. On another site a &quot;supplier reference&quot; field agreed 96% of the time. Neither column was called mpn. So we now test every code-like column against the customer&#39;s codes, whatever the column is called.\u003C/p>\n\u003Cp>Matching is still in progress, and we are not going to round it up. In our test environment the share of the customer&#39;s models with at least one competitor match went from about 20% to about 46% in one market, and from about 5% to about 38% in the other. The target is above 65%, and we are not there yet. Automatic matches were right at model level about 94% of the time, measured with the EAN hidden. For more on why matching is the hard half, see \u003Ca href=\"/blogs/product-matching-price-monitoring-2026\">product matching for price monitoring\u003C/a> and \u003Ca href=\"/blogs/ean-gtin-matching-failure-modes-2026\">why EAN matching fails\u003C/a>.\u003C/p>\n\u003Ch2 id=\"what-it-took\">What It Took\u003C/h2>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Step\u003C/th>\n\u003Cth>Time on this project\u003C/th>\n\u003C/tr>\n\u003C/thead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Intake: catalogue fill rates, competitor list, markets\u003C/td>\n\u003Ctd>about 2 hours\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Finding the JSON endpoints on every site\u003C/td>\n\u003Ctd>about half a day\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Building the scrapers with the same columns\u003C/td>\n\u003Ctd>about 1 day\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>First runs, one at a time, with a check after each\u003C/td>\n\u003Ctd>about 3 days of wall clock\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Cutting requests\u003C/td>\n\u003Ctd>about 35 minutes of study and 3 hours of changes\u003C/td>\n\u003C/tr>\n\u003C/tbody>\u003C/table>\n\u003Cp>The time is agent time. Most of the build was done by an AI agent working in our portal through the \u003Ca href=\"/blogs/build-web-scraper-with-claude-mcp-2026\">Scrapewise MCP server\u003C/a>, with a person deciding the open questions.\u003C/p>\n\u003Cp>Next time we will decide the cheapest daily source and the code columns before building, not after the first runs. Most of the rework came from doing that step late.\u003C/p>\n\u003Ch2 id=\"how-the-same-setup-scales\">How the Same Setup Scales\u003C/h2>\n\u003Cp>Nothing in this setup is tied to the number of competitors. Each new site is one more daily price scraper, one more one-time enrichment run and one more matching config, all in the same groups with the same column names. The same pattern works for any number of competitor sites and tens of thousands of links.\u003C/p>\n\u003Cp>The platform limits are sized for that. A link list fed automatically from another scraper can hold up to about 700,000 links. A bulk URL upload takes 5,000 URLs in one JSON call, or up to 100,000 as JSON Lines. A customer catalogue can come in as a CSV or JSON Lines file of up to 200 MB, or an Excel file of up to 36 MB. And a run has no time limit: it stops when the list is done, the wallet is empty, or you stop it.\u003C/p>\n\u003Ch2 id=\"if-you-want-the-same-setup\">If You Want the Same Setup\u003C/h2>\n\u003Cp>You can build this yourself in the portal, or have an agent build it for you through MCP. The parts that did the work are all standard: upload-file scrapers for your catalogue, API and HTML scrapers with link lists, after-scrape rules, a sample scrape before each full run, and the matching tab. The data comes out through the REST API or a CSV/Excel export.\u003C/p>\n\u003Cp>If you would rather hand it over, \u003Ca href=\"/use-cases/competitor-price-tracking\">competitor price tracking\u003C/a> is the managed version of the same thing. Every account starts with 5 free requests.\u003C/p>\n",{"title":11,"description":12,"badge":13,"benefits":14},"Frequently asked questions","Competitor price monitoring case study: questions answered","FAQ",[15,18,21,24,27],{"title":16,"description":17},"How did you cut daily requests by 88%?","We used each site's largest page and batch sizes, removed scrapers that read the same pages twice, switched one site from a browser to a plain fetch, and moved data that never changes, such as EANs and images, into one-time runs. On the sites where we changed the daily source, that took daily requests from 53,717 to 6,302.",{"title":19,"description":20},"Why use a site's API instead of scraping the HTML?","The JSON a shop's own pages load survives redesigns, carries exact ids and prices, and usually returns many products per request. HTML selectors break when the theme changes, so we use them only when no structured source exists.",{"title":22,"description":23},"What is a one-time enrichment run?","It is a single item-level scrape of data that does not change daily, such as EAN, maker codes, size, colour, brand and image. The daily price rows are joined to it on a stable product key, so the daily run only has to fetch prices.",{"title":25,"description":26},"Should I monitor price per size or per model?","Test it before you decide. We counted how often all sizes of a model had the same price and found 80% to 94% depending on the site, never 95%. So we kept per-size prices where they cost no extra request.",{"title":28,"description":29},"How long did the setup take?","About four days of agent time for intake, finding endpoints, building the scrapers and the first runs. Most of that went into rework, which is why we now choose the cheapest daily source before building.","Competitor Price Monitoring Case Study: 88% Fewer Requests","How we set up daily competitor price monitoring for a retailer in two EU markets: API first, HTML last, and 88% fewer requests per day.","Raivo Kartau",null,"Pricing",9,[37,42,48],{"slug":38,"title":39,"image":40,"date":5,"category":34,"excerpt":41},"price-per-unit-comparison-pack-sizes-2026","Price per Unit: How to Compare Competitor Prices Across Pack Sizes","/img/news/price-per-unit-comparison-pack-sizes-2026.png","A yarn retailer's competitors sell 25 g, 50 g, 100 g and 226 g balls. How we turned their competitor links into one comparable price per 50 g.",{"slug":43,"title":44,"image":45,"date":46,"category":34,"excerpt":47},"price2spy-vs-wiser-price-monitoring-2026","Price2Spy vs Wiser: Which Price Monitoring Tool Should You Choose in 2026?","/img/news/price2spy-vs-wiser-price-monitoring-2026.png","24 Sep 2026","Price2Spy vs Wiser in 2026: which has the more comprehensive feature set, why refresh speed is worthless without match accuracy, and the 50-SKU test to run.",{"slug":49,"title":50,"image":51,"date":46,"category":34,"excerpt":52},"prisync-vs-competera-price-monitoring-2026","Prisync vs Competera: Which Pricing Tool Should You Choose in 2026?","/img/news/prisync-vs-competera-price-monitoring-2026.png","Prisync vs Competera in 2026: $99–$399/month published plans against a quoted enterprise contract, and the one question that decides it — rules or recommendations.",{"slug":54,"title":55},"find-hidden-json-api-shop-page-2026","How to Find the Hidden JSON API Behind a Shop Page (and Why It Is Cheaper)",{"slug":57,"title":58},"build-web-scraper-with-claude-mcp-2026","Build a Web Scraper with Claude and MCP: A Step-by-Step Walkthrough",[60,64,67,70,73,76,79,82,85],{"level":61,"text":62,"id":63},2,"The Setup in One Picture","the-setup-in-one-picture",{"level":61,"text":65,"id":66},"API First, HTML Last","api-first-html-last",{"level":61,"text":68,"id":69},"How the Daily Requests Fell by 88%","how-the-daily-requests-fell-by-88",{"level":61,"text":71,"id":72},"The Mistakes That Cost Us Most","the-mistakes-that-cost-us-most",{"level":61,"text":74,"id":75},"Same Column Names on Every Scraper","same-column-names-on-every-scraper",{"level":61,"text":77,"id":78},"Matching: Codes First, Text Last","matching-codes-first-text-last",{"level":61,"text":80,"id":81},"What It Took","what-it-took",{"level":61,"text":83,"id":84},"How the Same Setup Scales","how-the-same-setup-scales",{"level":61,"text":86,"id":87},"If You Want the Same Setup","if-you-want-the-same-setup",[],1790769637104]