Match the same product across different sites· Lesson 6 of 6

Measure your match rate honestly

The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.

  • 11 min read
  • No account needed

"We match 94% of products" is a sentence with no meaning attached until somebody says what the denominator was. It is also the single most common way both vendors and internal teams mislead themselves — usually without intending to, because the flattering denominator is the one that is easiest to compute.

Four denominators, four very different numbers

Same system, same day. Only the question changed.

DenominatorWhat it answersHonest?
Products we attempted to matchHow often the matcher produced somethingNo — excludes everything you never tried
Products in your catalogueHow much of your range you have competitor data forYes, and usually the one that matters
Products you actively compete onCoverage where it has commercial consequenceYes, and often the most useful
Products that exist on both sidesMatcher quality in isolationYes, for engineering. Not for a business claim.

Two numbers, not one

Match rate alone cannot tell you whether a system is good, because it says nothing about whether the matches are right. Two measurements are needed and they trade against each other.

Precision: of the matches you published, what share are correct. This is the one that protects your credibility, and it is the one a pricing manager cares about, because a single visibly wrong comparison costs more trust than ten missing ones.

Recall: of the products that could have been matched, what share you found. This is coverage, and it is the one the person who bought the system asks about.

Pushing either to the limit destroys the other. A system that matches everything has terrible precision; a system that only matches barcodes has excellent precision and poor recall. The right balance comes from the cost asymmetry in lesson one, and for most commercial price monitoring that means favouring precision.

Building a sample you can trust

Two hundred hand-labelled pairs is enough to be useful and small enough that it actually gets done.

  1. 1

    Sample randomly, then stratify

    Random across the whole catalogue, then top it up so each major category and each band is represented. Sampling only from matches you already made measures nothing.

  2. 2

    Label by hand, with the pages open

    Open both listings and decide. Allow three answers — same, different, genuinely ambiguous — and keep the ambiguous ones, because their share is itself a finding.

  3. 3

    Have a second person label a slice

    Fifty pairs, independently. Where two careful humans disagree is the ceiling on what any automated system can achieve, and it is usually lower than people expect.

  4. 4

    Score the system against the labels

    Precision and recall separately, and broken down by category. The aggregate hides everything interesting.

  5. 5

    Hold the sample out

    Do not tune thresholds against the set you measure with. If you do, you are measuring how well you memorised two hundred pairs.

  6. 6

    Re-measure quarterly

    Both catalogues drift. A weighting tuned a year ago is describing a world that has moved.

Report it by category or you have reported nothing

An aggregate match rate averages a solved problem with an unsolved one and tells you about neither. Groceries and electronics carry barcodes and will pull the average up; furniture, fashion and bikes will pull it down, and those are often exactly the categories where someone is waiting on an answer.

A broken-down table — category, products, matched, precision — is immediately actionable. It shows where to spend effort, and it stops the conversation where someone quotes the aggregate at a category it does not describe.

Worked example: precision and recall from 200 labelled pairs

Two hundred pairs, labelled by a person, held back from everything used to tune the matcher. Of those, 120 are genuinely the same product and 80 are not, and the system proposed 110 matches. Both headline numbers come out of this one table and they move in opposite directions. Loosening the thresholds until recall reaches 95% would add about ten more true positives and take them from the borderline band, which is where the false ones live too — and for commercial pricing those 6 false positives already cost more than the 16 misses do.

System said matchSystem said no match
Genuinely the same product (120)104 — true positives16 — missed
Genuinely different (80)6 — false, and expensive74 — correctly rejected
Precision: 104 ÷ 110 = 94.5%Recall: 104 ÷ 120 = 86.7%

What usually goes wrong

Every item here produces a number that is higher than the truth and impossible to argue with.

  • Reporting one figure. Precision and recall trade against each other, so a single "accuracy" number can be improved by making the matcher either braver or more timid, and neither is information.
  • Labelling the sample from the system's own output. Pairs it never proposed cannot appear, so recall becomes unmeasurable by construction and comes out looking perfect.
  • Tuning on the evaluation set. The resulting number measures memorisation. Hold the 200 pairs back and do not look at them while adjusting anything.
  • Reporting one rate across all categories. Electronics with barcodes and fashion without them have completely different ceilings, and the blended figure describes neither of them.
  • Quietly changing the denominator when the number disappoints. Narrowing the scope explicitly and saying so is honest, and improving the input data is honest. The third option is not.
  • Not establishing the human ceiling. Have a second person label a slice: if two people disagree on 8% of pairs, a matcher scoring 92% is at the limit of the exercise rather than failing it.

What to do with a number you do not like

Sometimes the honest measurement is bad. We have seen 18% on a bicycle catalogue — a genuinely hard category where the same frame is sold under different model names with different component specs, and the number was a correct description of the problem rather than a broken system.

There are three defensible responses and one that is not. You can narrow the scope, to the products where matching is reliable and the commercial stakes are real. You can change the input, by sourcing identifiers from the manufacturer rather than inferring them from retailer titles. You can accept it and report it, with the uncovered products clearly labelled.

What you cannot do is quietly change the denominator, which is the response the incentives push hardest towards. A match rate that improved because the question got easier is not an improvement, and the person it fools first is you.

The general rule, and it is the same one as everywhere else in this course: publish what you measured. A page, a report or a feed that says plainly where the data does not reach is worth more than one that pads the gap — because the first can be acted on, and the second will be believed.

One question remains, and it is the one legal teams ask first. The final course covers what you are allowed to collect and how to answer for it.

Course 6: the legal and ethical side

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.