[{"data":1,"prerenderedAt":238},["ShallowReactive",2],{"learn-lesson-matching-products-across-sites-when-there-is-no-barcode":3},{"course":4,"lesson":67,"index":205,"outline":206,"prev":236,"next":237},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"matching-products-across-sites",5,"You have data from more than one site","6 lessons, about 70 minutes","Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Matching Products Across Sites: A Free 6-Lesson Course","How to match the same product across different retailers. GTIN and EAN matching, fuzzy title matching, variants and multipacks, confidence scoring, review queues, and measuring match rate honestly. Free, ungated.","product matching, product data matching, gtin matching, ean barcode matching, fuzzy product matching, sku mapping, competitor price matching, product catalogue matching","A free course on matching the same product across different retailers","Six written lessons on product matching: identifiers first, fuzzy matching when there is no barcode, variants and multipacks, confidence scores and review queues.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course five","Match the same product across different sites","A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.",{"label":21,"url":22},"Start with lesson one","/learn/matching-products-across-sites/why-matching-is-the-hard-part",{"label":24,"url":25},"See the data a run returns","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Explain why matching, not collection, is where price monitoring projects fail","Use GTIN, EAN, UPC and MPN correctly, including the checks that stop a bad identifier poisoning a match","Build a defensible fuzzy match for the majority of products that carry no usable barcode","Handle variants, bundles and multipacks without silently comparing a three-pack to a single unit","Attach a confidence score to every match and route the uncertain ones to a human","Measure your match rate with a denominator that tells the truth",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Pricing and e-commerce teams who have competitor data and cannot trust the comparisons","Developers building the matching layer and looking for the failure modes before they find them in production","Anyone evaluating a price monitoring vendor who wants to ask better questions than \"how many sites do you cover\"","Not written for",[44,45],"Teams who do not yet have data from more than one source — collection comes first, see course one or three","Anyone expecting a machine learning tutorial. The techniques here are mostly deterministic, because deterministic matches are the ones you can defend to a pricing manager.",{"title":47,"intro":48},"The six lessons","Lessons two and three are the mechanics. Lesson four is the one that catches most teams out, and lesson six is the one that stops you believing your own numbers.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions that come up as soon as somebody looks closely at a match table.",[54,57,60,63],{"title":55,"description":56},"What match rate should I expect?","It depends almost entirely on the category, and anyone who quotes you a single number across all categories is not being careful. Electronics and groceries carry barcodes and match well. Fashion, furniture and bikes are hard because the same frame is sold under different model names with different component specs. We have seen a real project stall at 18% on bicycles, which was a correct measurement of a genuinely hard category rather than a broken system.",{"title":58,"description":59},"Can I just match on product title?","You can, and you will get a result that looks plausible and is wrong often enough to be dangerous. Lesson three is about doing it properly when there is no alternative, which there frequently is not. The important part is not the string algorithm — it is attaching a confidence score and refusing to auto-publish the weak ones.",{"title":61,"description":62},"Should a human be in the loop?","Yes, for the middle band. The strong matches do not need a person and the hopeless ones do not deserve one. The value of a review queue is entirely in how well you have sorted it, which is lesson five.",{"title":64,"description":65},"Is this specific to Scrapewise?","No. Lesson five mentions how we structure a review queue, and that lesson is labelled. The rest is method and applies to a spreadsheet, a Python script or any vendor's matching engine.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":196,"next_step":201},"when-there-is-no-barcode","No barcode","Matching when there is no barcode","Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.","12 min",false,{"title":75,"description":76,"keywords":77},"Fuzzy Product Matching Without a Barcode","How to match products on title, brand and attributes when no identifier exists. Normalisation, blocking, multi-signal scoring and the mistakes that make fuzzy matching unreliable.","fuzzy product matching, string similarity matching, product title matching, brand normalisation, record linkage, entity resolution ecommerce",[79,84,100,105,111,143,148,176,187],{"type":80,"paragraphs":81},"prose",[82,83],"Most of your catalogue will not have a usable identifier on both sides. That is the normal case, not a data quality failure, and the work here is to build something that is defensibly right rather than something that feels clever.","The structure is four steps: normalise so the comparison is fair, block so you are not comparing everything to everything, score on several signals rather than one, and threshold into three bands instead of two.",{"type":85,"title":86,"items":87},"steps","The four steps",[88,91,94,97],{"title":89,"text":90},"Normalise both sides identically","Lowercase, strip punctuation, collapse whitespace, unify units (1l and 1000ml, 1kg and 1000g), remove retailer noise words like \"free delivery\" and \"new\". The single highest-value step is a brand alias table — Bosch Blue and Bosch Professional are the same brand, and no string algorithm will ever work that out.",{"title":92,"text":93},"Block before you compare","Comparing fifty thousand products against fifty thousand is 2.5 billion comparisons. Blocking cuts that by only comparing within a shared key — same brand, or same category, or sharing a rare token. You lose a small number of cross-block matches and gain a system that finishes.",{"title":95,"text":96},"Score on several signals, not one","Title similarity, brand agreement, numeric attribute agreement, pack size agreement and price ratio. Each contributes; none decides alone. The multi-signal part is what makes it defensible.",{"title":98,"text":99},"Threshold into three bands","Auto-accept, review, auto-reject. Two bands is the mistake — it forces every borderline case into a wrong answer.",{"type":101,"variant":102,"title":103,"text":104},"callout","note","The algorithm matters less than the normalisation","Teams spend days choosing between Levenshtein, Jaro-Winkler, token-set ratio and embeddings. In practice the difference between them is small compared to the difference made by a brand alias table and consistent unit handling. Normalise properly with a mediocre algorithm and you will beat a sophisticated algorithm on raw strings.",{"type":80,"title":106,"paragraphs":107},"Numbers in titles are the strongest weak signal",[108,109,110],"Product titles are mostly marketing words, which are the least discriminating part of the string. The numbers are where the information is: 18V, 55Nm, 4.0Ah, 256GB, 1.5L.","A practical rule that outperforms its simplicity: extract every number-with-unit from both titles and treat disagreement as near-fatal. Two titles that are 90% similar as strings but say 128GB and 256GB are different products, full stop, and string similarity will never tell you that because one character differs out of forty.","The same logic applies in reverse. Two titles that share little text but agree on every numeric attribute, the brand, and sit within a few percent on price are very probably the same thing described by two different marketing departments.",{"type":112,"title":113,"intro":114,"headers":115,"rows":119},"table","A workable signal weighting","A starting point, not a law. Tune on a hand-labelled sample — see lesson six.",[116,117,118],"Signal","Weight","Notes",[120,124,127,131,135,139],[121,122,123],"Brand agreement","Gate","Different brands should not match at all. Treat as a filter, not a score.",[125,122,126],"Numeric attribute agreement","Any disagreement on a spec number is near-fatal",[128,129,130],"Title token overlap","40%","After normalisation, ignoring stopwords",[132,133,134],"Model or part number match","35%","Huge when present, absent on most listings",[136,137,138],"Pack size agreement","15%","Covered properly in lesson four",[140,141,142],"Price plausibility","10%","Weak signal, real one. A 4x gap on a supposed match is a warning.",{"type":80,"title":144,"paragraphs":145},"Why price is a signal and also a trap",[146,147],"Price belongs in the score because an implausible ratio is genuine evidence against a match. It gets a small weight for a reason, though: you are building this system to detect price differences, so weighting price heavily means the system becomes least confident exactly when a competitor does something interesting.","Use it as a tiebreaker and a warning flag, never as a primary. The rule we settled on is that a price ratio beyond about three to one demotes a match to review regardless of how well everything else scored — because at that ratio, body-only versus full-kit is a more likely explanation than a genuine undercut.",{"type":112,"title":149,"intro":150,"headers":151,"rows":155},"Worked example: why blocking is not optional","Your 1,200 products against one competitor's 29,000 listings. Comparing every pair is nearly thirty-five million comparisons for one competitor on one run, and the great majority of them are a kettle against a garden hose. Each divisor below is just the number of distinct values in that field, so the arithmetic carries over to your own catalogue with your own figures.",[152,153,154],"Approach","Candidate pairs","What it removes",[156,160,164,168,172],[157,158,159],"Every product against every listing","1,200 × 29,000 = 34,800,000","nothing",[161,162,163],"Block on brand, with 60 brands in play","÷ 60 = about 580,000","every pair whose brands disagree",[165,166,167],"Block on brand and category, with 6 categories","÷ 6 = about 96,000","the kettle against the garden hose",[169,170,171],"Also require a shared number in the title, roughly 1 pair in 13","÷ 13 = about 7,400","most of the remaining near-misses",[173,174,175],"Of those, how many a person ever sees","a few hundred","which is the whole reason for three bands rather than one threshold",{"type":177,"title":178,"intro":179,"items":180},"list","What usually goes wrong","Separate from the scoring mistakes above, these are the ones that come from how the matcher is run and evaluated.",[181,182,183,184,185,186],"Skipping the blocking step and reaching for a faster machine. Thirty-five million comparisons is a design problem, not a performance problem, and the index that fixes it is an afternoon's work.","Blocking on a field that is not always populated. A block on category silently discards every listing where the category was not extracted, and those pairs never get the chance to match at all.","Evaluating on the pairs the system found. The pairs it missed are invisible by construction, so recall can only be measured against a hand-labelled sample.","Running the matcher over every row every night. Confirmed pairs are facts; re-deriving them nightly means the answer can change without anyone having changed anything.","Having no band for \"do not know\". Forcing every pair to match or not match converts an honest uncertainty into a wrong number on a dashboard.","Reading the score as a probability. It is a weighted sum whose weights you chose: 0.87 is an ordering, not a seven-in-eight chance of being right.",{"type":177,"title":188,"items":189},"Mistakes that make fuzzy matching untrustworthy",[190,191,192,193,194,195],"One threshold instead of three bands, which forces a wrong answer on every borderline pair","No brand gate, so two different manufacturers' similar products match on generic words","Treating a 95% string score as certainty when the 5% is the capacity or the pack size","Normalising one side and not the other, which quietly halves every score","Tuning the thresholds against the same examples you built them from, which measures memorisation rather than accuracy","Never re-checking. A weighting tuned on last year's catalogue drifts as both catalogues change.",[197,198,199,200],"Normalise, block, score on multiple signals, threshold into three bands.","A brand alias table and consistent unit handling beat a better string algorithm.","Numbers in titles carry the information. Treat spec-number disagreement as near-fatal.","Keep price as a low-weight tiebreaker — weighting it highly makes you least confident exactly when something interesting happens.",{"text":202,"label":203,"url":204},"Next: the failure that survives every technique above — packs, bundles and variants.","Lesson 4: variants, bundles and multipacks","/learn/matching-products-across-sites/variants-bundles-and-multipacks",2,[207,213,219,220,225,231],{"slug":208,"navTitle":209,"title":210,"summary":211,"time":212,"needsAccount":73},"why-matching-is-the-hard-part","Why matching is hard","Why matching is the hard part","The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.","10 min",{"slug":214,"navTitle":215,"title":216,"summary":217,"time":218,"needsAccount":73},"identifiers-first-gtin-ean-mpn","Identifiers first","Identifiers first: GTIN, EAN, UPC and MPN","What each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.","11 min",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":221,"navTitle":222,"title":223,"summary":224,"time":72,"needsAccount":73},"variants-bundles-and-multipacks","Variants and packs","Variants, bundles and multipacks","The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.",{"slug":226,"navTitle":227,"title":228,"summary":229,"time":218,"needsAccount":230},"score-confidence-and-build-a-review-queue","Confidence and review","Confidence scores and a review queue worth using","Why one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.",true,{"slug":232,"navTitle":233,"title":234,"summary":235,"time":218,"needsAccount":73},"measure-your-match-rate-honestly","Measure it honestly","Measure your match rate honestly","The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.",{"slug":214,"navTitle":215,"title":216,"summary":217,"time":218,"needsAccount":73},{"slug":221,"navTitle":222,"title":223,"summary":224,"time":72,"needsAccount":73},1791047867332]