Match the same product across different sites· Lesson 1 of 6

Why matching is the hard part

The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.

  • 10 min read
  • No account needed

Teams start price monitoring projects worrying about collection. Can we read the pages, will we get blocked, how fresh will it be. Those are solvable and the answers are mostly engineering.

Then the data arrives and the real problem shows up: you have forty thousand rows from six retailers and no reliable way to say which of them refer to the same object. Every comparison, every price index, every alert depends on that mapping being right, and nothing in the collection layer helps you build it.

The differences are deliberate

It is tempting to treat the variation in how retailers describe products as sloppiness. It is not. A retailer's product title is a marketing asset — it is tuned for their own search, their own SEO, their own customer. Two shops selling the identical item have active reasons to describe it differently, and a retailer with an exclusive-looking model name has a reason to make comparison harder.

So you are not cleaning up an accident. You are reconciling six catalogues that were each built for a different purpose, by people with no interest in them lining up.

One product, six retailers

A composite built from patterns we see constantly. Every column is a correct description of the same physical object.

RetailerTitle as publishedWhat makes it hard
ABosch Professional GSB 18V-55 Combi Drill (2x4.0Ah)Battery spec in the title, pack contents implied
BGSB18V-55 Cordless Hammer Drill Driver KitNo brand word, model number unspaced, different category noun
CBosch Blue 18V Combi Drill + 2 Batteries + CaseSub-brand instead of model, contents as a list
DBosch GSB 18V-55 (Body Only)Looks like the same product. It is not — no batteries.
EAkkuschrauber Bosch GSB 18V-55 ProfessionalDifferent language, different word order
FBosch 06019H5302Manufacturer part number as the entire title

What a wrong match costs, specifically

A missed match costs you visibility. You do not see a competitor's price, your coverage is lower than you thought, and the index is computed over a smaller set. That is a real cost and it is recoverable.

A false match costs you money and credibility. It produces a confident, precise, wrong number, and that number goes into a repricing rule or a board slide. The first time a pricing manager finds one, they stop trusting the whole feed — and they are right to, because they have no way to know which other rows are wrong.

This asymmetry should drive the entire design. Being conservative and leaving a product unmatched is nearly always cheaper than guessing. Systems tuned to maximise match rate are tuned for the wrong thing.

The five ways the same product diverges

Worth internalising, because each one needs a different handling strategy.

  • Naming — brand word present or absent, model number spaced or not, category noun varying
  • Language and locale — the same product in a different language, with different units and decimal separators
  • Pack and bundle — single unit, multipack, with-accessories, body-only
  • Variant — colour, size, capacity, which may or may not affect price
  • Identifier availability — some retailers publish a barcode, most do not, and some publish the wrong one

Worked example: what one false match costs

A single wrong pair, followed by an automatic repricing rule, over thirty days. Their listing is the 1-litre bottle and yours is the 2-litre; the two titles agree on brand, product and very nearly every word, which is why this match scored higher than most of your correct ones. You sell 40 units a day at €18.00 with €5.40 of margin, so the product costs you €12.60.

Before the matchAfter the rule acts
Their price, as matched—€9.95, being the 1-litre bottle
Your price€18.00€9.85
Margin per unit€5.40−€2.75
Units in thirty days1,2001,200, and probably more
Margin over thirty days€6,480−€3,300
Swing€9,780, on one SKU

What usually goes wrong

Matching goes wrong in the design rather than in the code, and nearly always in the same direction.

  • Tuning for match rate. It is the metric that rewards exactly the behaviour above, because the pairs it adds last are the ones most likely to be wrong.
  • Treating a missed match and a false match as the same size of error. One costs you visibility on a line; the other costs margin, quietly and continuously.
  • Assuming the collection layer helps. A flawless scrape of a wrong pair is a precisely wrong number, which is considerably harder to notice than a missing one.
  • Expecting retailers' titles to converge. They are marketing assets with a search term inside them, written differently on purpose, and no amount of cleaning turns them into a shared key.
  • Deferring the question until the data has arrived. Matching decides whether the feed is intelligence or fiction, so it belongs in the plan rather than in the clean-up.

The order of operations for the rest of the course

Identifiers first, because when they exist they are nearly free and nearly certain. Then fuzzy matching for everything else, with a confidence score attached rather than a yes or no. Then variant and pack handling, which is the step that catches row D. Then a review queue for the band in the middle. Then measurement, so you know what you actually have rather than what you hope you have.

Skipping straight to fuzzy matching is the common mistake, and it means doing hard probabilistic work on products where a barcode would have answered the question outright.

Next: the cheap certain wins — barcodes, part numbers, and the checks that keep them honest.

Lesson 2: identifiers first — GTIN, EAN, MPN

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.