A scraper needs addresses. Before anything can read a price, something has to produce a list of competitor product pages — ideally the complete list, ideally kept up to date as they add and drop lines.
This is the step teams most often do by hand, and it is the step where doing it by hand hurts most: a thousand URLs copied from a browser is a week of somebody's life, and it is stale by the time it is finished. There are four ways to avoid that, and they are worth trying in this order.
The four methods, easiest first
- 1
1. The XML sitemap
Most e-commerce platforms publish one, because Google needs it. Try /sitemap.xml and /robots.txt on the competitor's domain; robots.txt usually names the sitemap even when the default path is wrong. What you get is often a sitemap index pointing at a dozen child files, one of which is products. This is the best possible outcome: a complete, maintained, machine-readable list of every product page, published deliberately. On a well-run store it takes about five minutes.
- 2
2. Category pages
When there is no product sitemap, walk the category tree instead: open a listing page, take every product link on it, follow the pagination, repeat. Slower and noisier than a sitemap — you will pick up promotional tiles and cross-sell blocks — but it works on nearly every store, and it has one genuine advantage: it tells you which category the competitor files a product under, which is useful context for matching later.
- 3
3. Their own site search
If you only care about a few hundred known products, searching the competitor's site for each EAN or manufacturer part number is often the fastest route, and it gives you something the other methods do not: a direct, high-confidence link between their URL and your SKU. That link is worth a great deal in lesson five. The limitation is obvious — it only finds what you already know to look for.
- 4
4. A URL extractor
A tool that takes the domain and returns the product URLs, handling the sitemap-or-crawl decision for you. We publish a free one; so do others. This is method one and two with the plumbing hidden, which is worth it when you are doing this for eight competitors rather than one.
Cleaning the list
Whatever method produced it, the raw list is never the list you want. Three passes fix most of it.
First, strip tracking parameters. A URL ending in ?utm_source=newsletter is the same page as the one without it, but a naive pipeline will treat them as two pages and bill you for both. Cut everything after the question mark unless you know a parameter is load-bearing — on some stores the variant selector lives there, and you will see it because the page genuinely changes when you remove it.
Second, collapse variants. A t-shirt in six sizes is often six URLs with the same price. Decide deliberately whether you are tracking the product or the variant, because tracking all six multiplies your bill by six and usually tells you one thing. Where sizes really are priced differently — shoes, tyres, anything sold by capacity — you do want them separately, and you should say so explicitly rather than letting the extractor decide.
Third, drop what is not a product. Category pages, blog posts, gift cards, bundles, store locators. They sneak in from every method and each one is a page you pay to read and a row that will never match.
Choosing a method per competitor
In practice you will use different methods for different sites, and that is fine.
| Situation | Use | Watch out for |
|---|---|---|
| Product sitemap exists and is current | Sitemap | Sitemaps that include discontinued lines for months |
| No sitemap, normal server-rendered listings | Category crawl | Infinite scroll with no page numbers |
| You only track 200 known EANs | Their site search | Zero-result pages returning a 200 status |
| Eight competitors, limited patience | URL extractor | JavaScript-only storefronts returning nothing |
| Login required to see prices | None of these | Do not. A price behind a login is not a public price. |
Worked example: one competitor's sitemap to a payable URL list
Start to finish on a single competitor, with the figure each step removes. The absolute counts will differ for your sites; the shape of the reduction will not.
- 1
Fetch /sitemap.xml and follow the index
Most sitemaps are an index pointing at several child files. You want only the product ones — if they are named products-1.xml through products-4.xml, take all four and none of the blog, category or store-locator files. Say the four together list 41,000 URLs.
- 2
Strip query strings and fragments, then deduplicate
Remove everything after ? and #. Tracking parameters, sort orders and session ids turn one page into several, and every duplicate is a page you pay to fetch twice. Say this leaves 38,600 distinct URLs.
- 3
Keep only the paths that match the product pattern
Open three product pages by hand and find the shared segment — /p/, /product/, /dp/, or a trailing numeric id. Everything that does not match is a category, a filter permutation or an article. Say 29,000 survive.
- 4
Decide products or variants, and write the decision down
If the sitemap lists every colour and size separately, those 29,000 URLs may be 9,000 products. Collapsing to the parent cuts the fetch count by roughly two-thirds and loses per-variant stock. Collapsing is usually right for pricing and usually wrong for availability, which is why the decision belongs in writing rather than in whoever built it.
- 5
Intersect with the SKUs you already chose
You picked 1,200 SKUs in lesson two. Only the overlap is worth fetching. If 700 of your 1,200 appear in their list, this competitor costs 700 pages a run, not 29,000 — and that single intersection is the difference between a project and a quote you walk away from.
What usually goes wrong
Discovery is the cheapest stage to get right and the most expensive to get wrong, because the error multiplies by every future run.
- Taking the whole sitemap. It is the single most expensive mistake in this lesson: 29,000 pages a run instead of 700, every run, forever.
- Missing the sitemap index and reading only the first child file, then wondering why a third of the catalogue never appears.
- Leaving tracking parameters attached. The same product arrives three times under three URLs, and the deduplication problem moves downstream into matching, where it is far harder to see.
- Discovering URLs once and never again. Competitors add and drop lines constantly, so a list built in January is measurably wrong by March and silently wrong the entire time.
- Pushing past a login wall because the data behind it is better. That is the one method on this list with a legal answer rather than a technical one, and course six is where it gets answered.
Keeping the list fresh
Catalogues move. A competitor adds forty lines before Christmas and drops a hundred in January, and a URL list captured once decays at maybe one to three percent a month depending on the category. The symptom is subtle: your page count stays flat while the number of pages that return a real price slowly falls, and your coverage quietly erodes.
The fix is boring and effective: re-run whichever method you used on a schedule — monthly is usually enough — and diff the result against the current list. New URLs get added, URLs that have disappeared from the sitemap for two consecutive runs get retired. That diff is also a small, genuinely interesting competitive signal in its own right: it is a list of what your rival started and stopped selling.
With addresses in hand, the next question is what to read off each page — and which tempting fields are a trap.
Lesson 4: extracting price, stock and shipping