In September 2026 we built two competitor price setups back to back. One covered full competitor catalogues in two EU markets, matched against a catalogue of about 100,000 SKUs, and cut daily requests from 53,717 to 6,302. The other compared a yarn retailer's competitor links per 50 g, with the price per unit now right on 100% of rows.
Both work now. Both took longer than they should have, and nearly all of the extra time went into fixing our own mistakes. This is the list, grouped by where they happened, with what each one cost and the check we run now.
None of these are exotic. That is the uncomfortable part.
Mistakes in Choosing the Source
1. Fetching every variant URL when they all return the same page
One shop has a URL for every size of every product. We fed them all to the scraper. Each returned the same product page, about 14 times per product, and the run produced no rows from them. That was about €4 of pages for nothing, which is cheap as mistakes go, and a day of confusion, which is not.
Check now: fetch three variant URLs of one product and compare the responses. If they are the same, use one URL per product.
2. Trusting a site search as a full catalogue
A site search looked like a cheaper way to list one shop's whole range. It returned 245,000 rows, but only about 91,800 different products. Deep search pages repeated the same tiles.
Check now: before switching to a new source, diff its product ids against the old source. Row counts are not enough.
3. Paying for a browser the site did not need
One site looked like it needed a rendered page with Super (residential proxy). A plain fetch gave the same rows. On our pricing that is €3.75 per 1,000 pages against €0.15.
Check now: run a plain fetch and a rendered fetch on the same pages and compare the rows before choosing the tier.
4. Using AI extraction where a JSON source existed
On the yarn project we started with AI extraction on every link. Most of those shops had a structured source: Shopify's variant JSON, the WooCommerce Store API, or microdata in the page. The structured sources were more accurate and cost about a ninth as much. We now use AI only on the under 10% of links where nothing structured has the pack weight. How we look for these sources is in finding the hidden JSON API behind a shop page.
Check now: look for a platform JSON endpoint, JSON-LD and microdata before building an AI or selector scraper.
Mistakes in Reading the Data
5. Reading JSON arrays by position
An endpoint returned sizes, SKUs and EANs as separate arrays. When one variant had no EAN, every EAN after it moved to the wrong size. Nothing failed. The EANs were just attached to the wrong rows.
Check now: prefer endpoints with one object per variant. Compare 5 rows against the live page, field by field.
6. Assuming the API price is the page price
In one market, the API price and the price on the page differed by a constant factor: the VAT difference between two countries. Some discounts also exist only on the page.
Check now: compare the API price with the page price on about 50 products per market. A constant factor becomes a fixed column. A page-only discount is recorded as a known limit, not hidden.
7. Asking AI to do arithmetic
We asked the extraction step for a price per 50 g. It was right on 18 of 69 rows, and only where the ball was 50 g, so the price was the answer. It never calculated anything. The story is in let the AI read and let rules do the maths.
Check now: the AI copies printed values only. Every calculation is an after-scrape rule. Recompute a sample by hand after the first run.
8. Treating a crossed-out price equal to today's price as a sale
Two shops show a "before" price that is the same as the price you pay. Our first version called those sales.
Check now: a sale price is filled only when it is lower than the regular price.
9. Reading stock text as a quantity
An accessory page said "40+ pcs in stock". The AI read that as the number of pieces in the pack, and the row ended up with two different units.
Check now: each accessory row carries exactly one unit, per 100 g or per piece, and the unit comes from the product, not from the stock line.
Mistakes With the Customer's Own Data
10. Cleaning the customer's Excel file by hand
A brand in the customer's catalogue is called "100%". Excel stores that as the number 1 with a percent format. Our cleaning script read the stored value, so about 540 brand cells per market turned into "1".
Check now: upload the customer's raw file as an upload-file scraper, which keeps the displayed text. If a cleaning step is unavoidable, compare displayed and stored values per column. Percentages, dates and codes with leading zeros are the usual victims.
11. Re-uploading a file already used in matching
Once a file is used in matching, a new upload of it only adds rows and fills empty cells, matched by a key column. We re-uploaded a corrected catalogue expecting a clean replacement.
Check now: get the file right before the first matching run. If it has to change later, plan it as a new upload with its own key column.
12. Different column names on different scrapers
A join read sku. One scraper had written variantSku. The join filled 0 rows and reported no error.
Check now: one fixed set of column names for every scraper in the project, decided before the first build: name, brand, sku, productId, variantId, size, colour, price, url, image, and so on.
Mistakes in How We Worked
13. Starting many large runs at once
We started 18 large first runs at the same time. It was not our best morning. Now first runs go one at a time, and each is checked before the next starts: rows against the site's total, how many rows have each column filled, duplicates on the product id.
Check now: one large run at a time, with a check after each.
14. Counting against our own list instead of the customer's
After the yarn build, the build notes counted more links as complete than an independent check with stricter rules did. One link in the customer's file had never been built, and two "sales" were the mistake in number 8.
Check now: count completeness against the customer's original file, not the builder's list, and have someone other than the builder do the count.
What We Changed in Our Process
Most of the rework on the multi-market project came from one ordering mistake. We decided the cheapest daily source and the code columns after the first runs, when they belong before the build. Moving that step forward is most of the time we expect to save next time.
The other change is smaller and it catches more. After every first run, we write one table row per scraper: rows, catalogue gap, fill per column, duplicates, verdict. Reading those rows takes a few minutes, and most of the mistakes above show up in them as an empty column or an odd count.
The full build that these came from is written up in the competitor price monitoring case study and the price per unit comparison. If you run your own scrapers, the status-code trap and self-healing scrapers cover the failures that come after the build. If you would rather someone else made these mistakes for you, competitor price tracking is the managed version.
Paste any URL — ScrapeWise handles the anti-bot
Managed infrastructure that adapts when sites change. No proxies, no code, no per-request fees.
Not ready to sign up? See 12 real Amazon rows, 972 columns →
97% accuracy on Amazon benchmarks · no credit card · book a 15-min call →
