A matching system that outputs yes or no is throwing away the most useful thing it knows, which is how sure it is. Once the score survives into the output, you can route by it: publish the top band, discard the bottom, and give a person the middle, which is usually a small fraction of the catalogue and the only part where human time pays for itself.
One score is not enough
This is a mistake we shipped and had to unpick. A single confidence number tends to conflate two separate questions: how certain are we that we read this price correctly, and how certain are we that these two products are the same thing.
They are independent. A price can be read perfectly from a page whose product is only probably the right match. A price can be a guess from a page that is unambiguously the right product. Collapsing those into one number produces outputs that are genuinely confusing to the person reading them — a high score on a row that is obviously wrong, and no way to tell which half was the problem.
Carry two: extraction confidence and match confidence. They route differently. Low extraction confidence is an engineering ticket. Low match confidence is a review item. Sending either to the wrong queue wastes the time of whoever receives it.
Three bands, by match confidence
Starting values. Lesson six is how to replace them with measured ones.
| Band | Score | Action | Typical share |
|---|---|---|---|
| Auto-accept | Valid identifier, or score above ~0.9 with all gates passed | Publish | Varies hugely by category |
| Review | ~0.6 to ~0.9, or any gate conflict | Human queue, ordered by value | Aim for under 10% of the catalogue |
| Reject | Below ~0.6, or a brand or spec gate failed | Discard, log the near-misses | The rest |
Order the queue by value, not by score
This is the difference between a review queue that works and one that quietly stops being opened.
Sorting by confidence puts your reviewer on the most ambiguous rows in the catalogue, which are also frequently the least important ones — obscure products nobody competes on. They spend an hour on hard calls with no business consequence, and the queue feels like a punishment.
Sort by potential impact instead: your sales volume for that product, multiplied by the price gap the match implies. Now the first twenty rows are the ones where being wrong actually costs something. An hour on that list is worth having, and the reviewer can see why.
What a review row needs on screen
The goal is a decision in under ten seconds. Anything that forces a tab switch breaks the rhythm.
- Both titles, with the differing tokens highlighted
- Both images, side by side — the fastest signal a human has
- Both prices, plus the implied ratio
- Both pack sizes and the key spec numbers, aligned
- Why the matcher was unsure, named explicitly: "pack size disagreement", "no shared model number"
- Three buttons: same, not the same, cannot tell. The third one is essential and routinely omitted.
Worked example: the same queue, ordered two ways
Five pairs waiting, one hour of somebody's attention, two orderings. Score ordering puts the reviewer on the pairs the system is least sure about; value ordering — monthly units times the price gap — puts them on the pairs where being wrong costs something. The two disagree completely: the score ordering spends the hour on €26.80 of exposure and the value ordering spends it on €7,860. Rows A and B never get looked at under the second ordering, and that is the correct outcome.
| Pair | Score | Monthly units | Price gap | At stake | Rank by score | Rank by value |
|---|---|---|---|---|---|---|
| A. Niche accessory | 0.52 | 4 | €1.20 | €4.80 | 1 | 5 |
| B. Discontinued line | 0.58 | 11 | €2.00 | €22.00 | 2 | 4 |
| C. Mid-range appliance | 0.71 | 90 | €14.00 | €1,260 | 3 | 2 |
| D. Top-selling TV | 0.74 | 300 | €22.00 | €6,600 | 4 | 1 |
| E. Spare part | 0.80 | 25 | €3.50 | €87.50 | 5 | 3 |
What usually goes wrong
A review queue is a product with exactly one user, and it fails for the reasons products fail.
- Carrying one score. Extraction confidence and match confidence answer different questions, and a blended number cannot tell you whether to re-scrape or to review.
- Setting thresholds from the shape of the histogram. The gap in the distribution is an artefact of your own weighting; the thresholds belong where the cost of a false match overtakes the cost of a missed one.
- Ordering by score. The hour goes on the cheapest pairs in the catalogue, exactly as in the table.
- A review row without the evidence on it. If the reviewer has to open two tabs to decide, throughput collapses and the queue stops being opened at all.
- Not persisting the decision. A queue that presents the same pair again next week is a queue that gets abandoned in week three.
- Keying the decision on a URL. Retailers change URLs, so key it on product identity — otherwise every confirmed match quietly expires the next time somebody tidies a slug.
Decisions are an asset — persist them
Every human judgement is expensive and permanently useful. Store it as an override keyed on the two product identities, not on row ids that change between runs, and make sure the next run reads it before it scores anything.
The failure mode here is brutal and common: a review queue that re-presents the same hundred pairs every week because the decisions were never written back. People stop opening it within a month, entirely reasonably.
Two refinements worth adding once the basics work. Record who decided and when, so a disputed comparison can be traced. And re-surface old decisions when the underlying titles change materially — a product that was correctly matched last year may have been replaced by a new generation under the same listing.
Finally: measuring what you have built, with a denominator that does not flatter you.
Lesson 6: measure your match rate honestly