Build a competitor price monitoring pipeline· Lesson 3 of 8

Finding every competitor product URL without copying them by hand

Four ways to get a competitor's full product URL list, ranked by how much work they are, and what to do when none of them work.

  • 12 min read

A scraper needs addresses. Before anything can read a price, something has to produce a list of competitor product pages — ideally the complete list, ideally kept up to date as they add and drop lines.

This is the step teams most often do by hand, and it is the step where doing it by hand hurts most: a thousand URLs copied from a browser is a week of somebody's life, and it is stale by the time it is finished. There are four ways to avoid that, and they are worth trying in this order.

The four methods, easiest first

  1. 1

    1. The XML sitemap

    Most e-commerce platforms publish one, because Google needs it. Try /sitemap.xml and /robots.txt on the competitor's domain; robots.txt usually names the sitemap even when the default path is wrong. What you get is often a sitemap index pointing at a dozen child files, one of which is products. This is the best possible outcome: a complete, maintained, machine-readable list of every product page, published deliberately. On a well-run store it takes about five minutes.

  2. 2

    2. Category pages

    When there is no product sitemap, walk the category tree instead: open a listing page, take every product link on it, follow the pagination, repeat. Slower and noisier than a sitemap — you will pick up promotional tiles and cross-sell blocks — but it works on nearly every store, and it has one genuine advantage: it tells you which category the competitor files a product under, which is useful context for matching later.

  3. 3

    3. Their own site search

    If you only care about a few hundred known products, searching the competitor's site for each EAN or manufacturer part number is often the fastest route, and it gives you something the other methods do not: a direct, high-confidence link between their URL and your SKU. That link is worth a great deal in lesson five. The limitation is obvious — it only finds what you already know to look for.

  4. 4

    4. A URL extractor

    A tool that takes the domain and returns the product URLs, handling the sitemap-or-crawl decision for you. We publish a free one; so do others. This is method one and two with the plumbing hidden, which is worth it when you are doing this for eight competitors rather than one.

Cleaning the list

Whatever method produced it, the raw list is never the list you want. Three passes fix most of it.

First, strip tracking parameters. A URL ending in ?utm_source=newsletter is the same page as the one without it, but a naive pipeline will treat them as two pages and bill you for both. Cut everything after the question mark unless you know a parameter is load-bearing — on some stores the variant selector lives there, and you will see it because the page genuinely changes when you remove it.

Second, collapse variants. A t-shirt in six sizes is often six URLs with the same price. Decide deliberately whether you are tracking the product or the variant, because tracking all six multiplies your bill by six and usually tells you one thing. Where sizes really are priced differently — shoes, tyres, anything sold by capacity — you do want them separately, and you should say so explicitly rather than letting the extractor decide.

Third, drop what is not a product. Category pages, blog posts, gift cards, bundles, store locators. They sneak in from every method and each one is a page you pay to read and a row that will never match.

Choosing a method per competitor

In practice you will use different methods for different sites, and that is fine.

SituationUseWatch out for
Product sitemap exists and is currentSitemapSitemaps that include discontinued lines for months
No sitemap, normal server-rendered listingsCategory crawlInfinite scroll with no page numbers
You only track 200 known EANsTheir site searchZero-result pages returning a 200 status
Eight competitors, limited patienceURL extractorJavaScript-only storefronts returning nothing
Login required to see pricesNone of theseDo not. A price behind a login is not a public price.

Worked example: one competitor's sitemap to a payable URL list

Start to finish on a single competitor, with the figure each step removes. The absolute counts will differ for your sites; the shape of the reduction will not.

  1. 1

    Fetch /sitemap.xml and follow the index

    Most sitemaps are an index pointing at several child files. You want only the product ones — if they are named products-1.xml through products-4.xml, take all four and none of the blog, category or store-locator files. Say the four together list 41,000 URLs.

  2. 2

    Strip query strings and fragments, then deduplicate

    Remove everything after ? and #. Tracking parameters, sort orders and session ids turn one page into several, and every duplicate is a page you pay to fetch twice. Say this leaves 38,600 distinct URLs.

  3. 3

    Keep only the paths that match the product pattern

    Open three product pages by hand and find the shared segment — /p/, /product/, /dp/, or a trailing numeric id. Everything that does not match is a category, a filter permutation or an article. Say 29,000 survive.

  4. 4

    Decide products or variants, and write the decision down

    If the sitemap lists every colour and size separately, those 29,000 URLs may be 9,000 products. Collapsing to the parent cuts the fetch count by roughly two-thirds and loses per-variant stock. Collapsing is usually right for pricing and usually wrong for availability, which is why the decision belongs in writing rather than in whoever built it.

  5. 5

    Intersect with the SKUs you already chose

    You picked 1,200 SKUs in lesson two. Only the overlap is worth fetching. If 700 of your 1,200 appear in their list, this competitor costs 700 pages a run, not 29,000 — and that single intersection is the difference between a project and a quote you walk away from.

What usually goes wrong

Discovery is the cheapest stage to get right and the most expensive to get wrong, because the error multiplies by every future run.

  • Taking the whole sitemap. It is the single most expensive mistake in this lesson: 29,000 pages a run instead of 700, every run, forever.
  • Missing the sitemap index and reading only the first child file, then wondering why a third of the catalogue never appears.
  • Leaving tracking parameters attached. The same product arrives three times under three URLs, and the deduplication problem moves downstream into matching, where it is far harder to see.
  • Discovering URLs once and never again. Competitors add and drop lines constantly, so a list built in January is measurably wrong by March and silently wrong the entire time.
  • Pushing past a login wall because the data behind it is better. That is the one method on this list with a legal answer rather than a technical one, and course six is where it gets answered.

Keeping the list fresh

Catalogues move. A competitor adds forty lines before Christmas and drops a hundred in January, and a URL list captured once decays at maybe one to three percent a month depending on the category. The symptom is subtle: your page count stays flat while the number of pages that return a real price slowly falls, and your coverage quietly erodes.

The fix is boring and effective: re-run whichever method you used on a schedule — monthly is usually enough — and diff the result against the current list. New URLs get added, URLs that have disappeared from the sitemap for two consecutive runs get retired. That diff is also a small, genuinely interesting competitive signal in its own right: it is a list of what your rival started and stopped selling.

With addresses in hand, the next question is what to read off each page — and which tempting fields are a trap.

Lesson 4: extracting price, stock and shipping

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.