Build a competitor price monitoring pipeline· Lesson 5 of 8

Matching competitor listings to your own catalogue

The stage that decides whether your feed is intelligence or fiction, and the denominator trick that makes bad match rates look good.

  • 13 min read
  • No account needed

This is the lesson that decides whether everything before it was worth doing. A competitor price is only meaningful attached to one of your products, and deciding that their listing and your SKU are the same sellable thing is harder than it sounds, permanently imperfect, and almost never discussed honestly in a sales demo.

It is also the stage where you should be most suspicious of numbers, including your own.

Four tiers of evidence

Matches are not binary. They sit on a ladder of confidence, and a well-run pipeline knows which rung each row is on.

Tier one is a shared global identifier: EAN, GTIN, UPC, ISBN. If both sides publish one and they agree, you are done — this is the only tier that is genuinely certain, and it is why lesson four told you to grab an EAN whenever the page shows one.

Tier two is a manufacturer part number plus brand. Strong, but MPNs get typed by hand into product feeds, so expect whitespace, hyphens and case to differ, and expect a small number of outright errors on both sides.

Tier three is attribute matching: brand, model, capacity, colour, pack size. This is where most real matching happens, and where most real errors live. Two litres against one litre, a twelve-pack against a single, last year's model number against this year's.

Tier four is fuzzy title similarity, usually with an embedding model. Useful for generating candidates a human then confirms. Not useful as a final answer, because it is confidently wrong in exactly the cases that cost the most money.

What a realistic match rate looks like

Measured against the honest denominator: products the competitor actually carries. These are the bands we see; your category decides where you land.

CategoryTypical achievableWhy
Books, media, games95%+Universal identifiers, published by everyone, rarely wrong
Branded electronics85-95%MPNs are usually present and usually correct
Grocery and FMCG70-90%EANs are common, but pack size and multipacks create genuine ambiguity
Fashion40-70%Season codes, colour names and exclusive colourways. Hard, and honestly hard.
Bikes, furniture, configurables30-60%Variants and bundles mean "the same product" often is not a well-defined idea
Own-brand0%There is nothing to match to. Exclude at scope time, per lesson two.

Fix your own data first

The instinct is to blame the competitor's messy listings. In our experience the larger share of the problem is on the home side, and it is cheaper to fix.

Missing EANs in your own catalogue are the biggest single cause of poor matching, and most retailers have more holes than they think — the field exists, it is populated for 60% of lines, and nobody has looked. Pack sizes buried inside a free-text title rather than held as a number are the second. Duplicate internal SKUs for the same physical product are the third: the competitor has one listing, you have three, and your match rate is immediately capped at a third for those lines.

Before investing in better matching, spend a day measuring: what proportion of your in-scope SKUs have a populated, valid EAN? That number predicts your ceiling better than any vendor's algorithm.

A workable process

  1. 1

    Match on identifiers first, and stop there where you can

    Run the EAN and MPN passes, mark those rows as high confidence, and take them out of the pool. Depending on category this is anywhere from a third to nearly all of your matches, and it needs no review.

  2. 2

    Generate candidates for the rest

    Use title and attribute similarity to propose, say, the three most likely counterparts for each unmatched product. The goal here is recall, not precision — you want the right answer to be in the list, not to be alone in it.

  3. 3

    Have a human confirm the candidates

    Someone who knows the category can confirm or reject a proposed match in a couple of seconds from a title, an image and a price. A few hundred of those is an afternoon, it is a one-off cost per product, and the decisions are reusable forever.

  4. 4

    Store the confirmed match, not the algorithm

    Once a human has said "their URL X is our SKU Y", that is a fact you own. Persist it. Re-deriving it on every run is how a pipeline that worked last month starts producing different answers this month for no visible reason.

  5. 5

    Review the unmatched pile monthly

    Unmatched is not a failure state, it is a queue. It also carries a signal worth reading: a sudden growth in unmatched rows for one competitor usually means they have relabelled a range or changed their title format, not that your matcher broke.

Decide what to do with a partial match

The genuinely difficult cases are not the failures, they are the near-misses. Their listing is the same model in a different colour. Theirs is a bundle with a case included. Theirs is the 2024 model and you stock the 2025.

There is no universally correct answer, but there is a correct process: decide the policy once, in writing, with the person who owns pricing — and record the reason a row was matched or excluded, so that six months later somebody can see why the 2024 model was being compared to the 2025 one. An undocumented matching policy becomes folklore within a quarter.

Worked example: the same match rate told three ways

A supplier reports 90%. The buyer counts 58.5%. Neither is lying; they are dividing by different numbers. Take 1,200 of your SKUs run against one competitor. The third figure is the one that belongs in a business case, because it is the only one you can reprice from without a person looking first.

StepCountWhat it gives you
SKUs you sent1,200the question you asked
Not stocked by this competitor at all420780 are even eligible to match
Eligible, but could not be resolved to a listing78702 matched
Of the 702, matched on a shared EAN430safe to act on unreviewed
Of the 702, matched on title and attributes only272needs a human before it moves a price
Vendor's headline: 702 ÷ 78090%true, and not the number you need
Yours: 702 ÷ 1,20058.5%coverage of the question you asked
Confident coverage: 430 ÷ 1,200about 36%what you can actually price from

What usually goes wrong

Matching is where a feed stops being data and starts being fiction, and it does so without raising anything.

  • Accepting a match rate with no denominator. Ask "out of what?" before the figure is written down anywhere, because once it is in a slide it will be quoted for a year.
  • Letting tier-three matches move prices. Title similarity puts a 2-litre bottle next to a 1-litre bottle and a 2026 model next to a 2024 one, and both look exactly like a competitor undercutting you by half.
  • Re-deriving matches on every run. A human confirmed that pair once; if the pipeline rediscovers it nightly it will eventually rediscover it differently, and nobody will be able to say when the number changed.
  • Fixing the matcher instead of your own data. Missing EANs and inconsistent pack sizes in your own catalogue are usually the largest single cause, and they are the only part of this you fully control.
  • Counting a variant match as a product match. Their listing is the 128 GB model and yours is the 256 GB: same brand, near-identical title, different thing, and the report says you are 180 euros overpriced.

A matched feed that silently loses a third of its rows is worse than no feed. Next: how to notice.

Lesson 6: scheduling and data quality

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.