Match the same product across different sites· Lesson 3 of 6

Matching when there is no barcode

Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.

  • 12 min read
  • No account needed

Most of your catalogue will not have a usable identifier on both sides. That is the normal case, not a data quality failure, and the work here is to build something that is defensibly right rather than something that feels clever.

The structure is four steps: normalise so the comparison is fair, block so you are not comparing everything to everything, score on several signals rather than one, and threshold into three bands instead of two.

The four steps

  1. 1

    Normalise both sides identically

    Lowercase, strip punctuation, collapse whitespace, unify units (1l and 1000ml, 1kg and 1000g), remove retailer noise words like "free delivery" and "new". The single highest-value step is a brand alias table — Bosch Blue and Bosch Professional are the same brand, and no string algorithm will ever work that out.

  2. 2

    Block before you compare

    Comparing fifty thousand products against fifty thousand is 2.5 billion comparisons. Blocking cuts that by only comparing within a shared key — same brand, or same category, or sharing a rare token. You lose a small number of cross-block matches and gain a system that finishes.

  3. 3

    Score on several signals, not one

    Title similarity, brand agreement, numeric attribute agreement, pack size agreement and price ratio. Each contributes; none decides alone. The multi-signal part is what makes it defensible.

  4. 4

    Threshold into three bands

    Auto-accept, review, auto-reject. Two bands is the mistake — it forces every borderline case into a wrong answer.

Numbers in titles are the strongest weak signal

Product titles are mostly marketing words, which are the least discriminating part of the string. The numbers are where the information is: 18V, 55Nm, 4.0Ah, 256GB, 1.5L.

A practical rule that outperforms its simplicity: extract every number-with-unit from both titles and treat disagreement as near-fatal. Two titles that are 90% similar as strings but say 128GB and 256GB are different products, full stop, and string similarity will never tell you that because one character differs out of forty.

The same logic applies in reverse. Two titles that share little text but agree on every numeric attribute, the brand, and sit within a few percent on price are very probably the same thing described by two different marketing departments.

A workable signal weighting

A starting point, not a law. Tune on a hand-labelled sample — see lesson six.

SignalWeightNotes
Brand agreementGateDifferent brands should not match at all. Treat as a filter, not a score.
Numeric attribute agreementGateAny disagreement on a spec number is near-fatal
Title token overlap40%After normalisation, ignoring stopwords
Model or part number match35%Huge when present, absent on most listings
Pack size agreement15%Covered properly in lesson four
Price plausibility10%Weak signal, real one. A 4x gap on a supposed match is a warning.

Why price is a signal and also a trap

Price belongs in the score because an implausible ratio is genuine evidence against a match. It gets a small weight for a reason, though: you are building this system to detect price differences, so weighting price heavily means the system becomes least confident exactly when a competitor does something interesting.

Use it as a tiebreaker and a warning flag, never as a primary. The rule we settled on is that a price ratio beyond about three to one demotes a match to review regardless of how well everything else scored — because at that ratio, body-only versus full-kit is a more likely explanation than a genuine undercut.

Worked example: why blocking is not optional

Your 1,200 products against one competitor's 29,000 listings. Comparing every pair is nearly thirty-five million comparisons for one competitor on one run, and the great majority of them are a kettle against a garden hose. Each divisor below is just the number of distinct values in that field, so the arithmetic carries over to your own catalogue with your own figures.

ApproachCandidate pairsWhat it removes
Every product against every listing1,200 × 29,000 = 34,800,000nothing
Block on brand, with 60 brands in play÷ 60 = about 580,000every pair whose brands disagree
Block on brand and category, with 6 categories÷ 6 = about 96,000the kettle against the garden hose
Also require a shared number in the title, roughly 1 pair in 13÷ 13 = about 7,400most of the remaining near-misses
Of those, how many a person ever seesa few hundredwhich is the whole reason for three bands rather than one threshold

What usually goes wrong

Separate from the scoring mistakes above, these are the ones that come from how the matcher is run and evaluated.

  • Skipping the blocking step and reaching for a faster machine. Thirty-five million comparisons is a design problem, not a performance problem, and the index that fixes it is an afternoon's work.
  • Blocking on a field that is not always populated. A block on category silently discards every listing where the category was not extracted, and those pairs never get the chance to match at all.
  • Evaluating on the pairs the system found. The pairs it missed are invisible by construction, so recall can only be measured against a hand-labelled sample.
  • Running the matcher over every row every night. Confirmed pairs are facts; re-deriving them nightly means the answer can change without anyone having changed anything.
  • Having no band for "do not know". Forcing every pair to match or not match converts an honest uncertainty into a wrong number on a dashboard.
  • Reading the score as a probability. It is a weighted sum whose weights you chose: 0.87 is an ordering, not a seven-in-eight chance of being right.

Mistakes that make fuzzy matching untrustworthy

  • One threshold instead of three bands, which forces a wrong answer on every borderline pair
  • No brand gate, so two different manufacturers' similar products match on generic words
  • Treating a 95% string score as certainty when the 5% is the capacity or the pack size
  • Normalising one side and not the other, which quietly halves every score
  • Tuning the thresholds against the same examples you built them from, which measures memorisation rather than accuracy
  • Never re-checking. A weighting tuned on last year's catalogue drifts as both catalogues change.

Next: the failure that survives every technique above — packs, bundles and variants.

Lesson 4: variants, bundles and multipacks

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.