Match the same product across different sites
A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.
What you will be able to do
- Explain why matching, not collection, is where price monitoring projects fail
- Use GTIN, EAN, UPC and MPN correctly, including the checks that stop a bad identifier poisoning a match
- Build a defensible fuzzy match for the majority of products that carry no usable barcode
- Handle variants, bundles and multipacks without silently comparing a three-pack to a single unit
- Attach a confidence score to every match and route the uncertain ones to a human
- Measure your match rate with a denominator that tells the truth
The six lessons
Lessons two and three are the mechanics. Lesson four is the one that catches most teams out, and lesson six is the one that stops you believing your own numbers.
- 01Why matching is the hard partThe same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.10 min read Read lesson 1 →
- 02Identifiers first: GTIN, EAN, UPC and MPNWhat each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.11 min read Read lesson 2 →
- 03Matching when there is no barcodeNormalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.12 min read Read lesson 3 →
- 04Variants, bundles and multipacksThe highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.12 min read Read lesson 4 →
- 05Confidence scores and a review queue worth usingNeeds an accountWhy one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.11 min read Read lesson 5 →
- 06Measure your match rate honestlyThe denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.11 min read Read lesson 6 →
Who this is for
Written for
- Pricing and e-commerce teams who have competitor data and cannot trust the comparisons
- Developers building the matching layer and looking for the failure modes before they find them in production
- Anyone evaluating a price monitoring vendor who wants to ask better questions than "how many sites do you cover"
Not written for
- Teams who do not yet have data from more than one source — collection comes first, see course one or three
- Anyone expecting a machine learning tutorial. The techniques here are mostly deterministic, because deterministic matches are the ones you can defend to a pricing manager.
Before you start
The questions that come up as soon as somebody looks closely at a match table.
It depends almost entirely on the category, and anyone who quotes you a single number across all categories is not being careful. Electronics and groceries carry barcodes and match well. Fashion, furniture and bikes are hard because the same frame is sold under different model names with different component specs. We have seen a real project stall at 18% on bicycles, which was a correct measurement of a genuinely hard category rather than a broken system.
Start at lesson one
The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.