[{"data":1,"prerenderedAt":226},["ShallowReactive",2],{"learn-lesson-matching-products-across-sites-measure-your-match-rate-honestly":3},{"course":4,"lesson":67,"index":6,"outline":194,"prev":224,"next":225},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"matching-products-across-sites",5,"You have data from more than one site","6 lessons, about 70 minutes","Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Matching Products Across Sites: A Free 6-Lesson Course","How to match the same product across different retailers. GTIN and EAN matching, fuzzy title matching, variants and multipacks, confidence scoring, review queues, and measuring match rate honestly. Free, ungated.","product matching, product data matching, gtin matching, ean barcode matching, fuzzy product matching, sku mapping, competitor price matching, product catalogue matching","A free course on matching the same product across different retailers","Six written lessons on product matching: identifiers first, fuzzy matching when there is no barcode, variants and multipacks, confidence scores and review queues.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course five","Match the same product across different sites","A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.",{"label":21,"url":22},"Start with lesson one","/learn/matching-products-across-sites/why-matching-is-the-hard-part",{"label":24,"url":25},"See the data a run returns","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Explain why matching, not collection, is where price monitoring projects fail","Use GTIN, EAN, UPC and MPN correctly, including the checks that stop a bad identifier poisoning a match","Build a defensible fuzzy match for the majority of products that carry no usable barcode","Handle variants, bundles and multipacks without silently comparing a three-pack to a single unit","Attach a confidence score to every match and route the uncertain ones to a human","Measure your match rate with a denominator that tells the truth",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Pricing and e-commerce teams who have competitor data and cannot trust the comparisons","Developers building the matching layer and looking for the failure modes before they find them in production","Anyone evaluating a price monitoring vendor who wants to ask better questions than \"how many sites do you cover\"","Not written for",[44,45],"Teams who do not yet have data from more than one source — collection comes first, see course one or three","Anyone expecting a machine learning tutorial. The techniques here are mostly deterministic, because deterministic matches are the ones you can defend to a pricing manager.",{"title":47,"intro":48},"The six lessons","Lessons two and three are the mechanics. Lesson four is the one that catches most teams out, and lesson six is the one that stops you believing your own numbers.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions that come up as soon as somebody looks closely at a match table.",[54,57,60,63],{"title":55,"description":56},"What match rate should I expect?","It depends almost entirely on the category, and anyone who quotes you a single number across all categories is not being careful. Electronics and groceries carry barcodes and match well. Fashion, furniture and bikes are hard because the same frame is sold under different model names with different component specs. We have seen a real project stall at 18% on bicycles, which was a correct measurement of a genuinely hard category rather than a broken system.",{"title":58,"description":59},"Can I just match on product title?","You can, and you will get a result that looks plausible and is wrong often enough to be dangerous. Lesson three is about doing it properly when there is no alternative, which there frequently is not. The important part is not the string algorithm — it is attaching a confidence score and refusing to auto-publish the weak ones.",{"title":61,"description":62},"Should a human be in the loop?","Yes, for the middle band. The strong matches do not need a person and the hopeless ones do not deserve one. The value of a review queue is entirely in how well you have sorted it, which is lesson five.",{"title":64,"description":65},"Is this specific to Scrapewise?","No. Lesson five mentions how we structure a review queue, and that lesson is labelled. The rest is method and applies to a spreadsheet, a Python script or any vendor's matching engine.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":185,"next_step":190},"measure-your-match-rate-honestly","Measure it honestly","Measure your match rate honestly","The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.","11 min",false,{"title":75,"description":76,"keywords":77},"How to Measure Product Match Rate Correctly","Why match rate is usually quoted against a flattering denominator, how to build a hand-labelled sample, measuring precision and recall separately, and reporting a low number well.","match rate measurement, precision and recall matching, product matching accuracy, price monitoring coverage, data quality metrics",[79,83,108,113,120,143,148,167,178],{"type":80,"paragraphs":81},"prose",[82],"\"We match 94% of products\" is a sentence with no meaning attached until somebody says what the denominator was. It is also the single most common way both vendors and internal teams mislead themselves — usually without intending to, because the flattering denominator is the one that is easiest to compute.",{"type":84,"title":85,"intro":86,"headers":87,"rows":91},"table","Four denominators, four very different numbers","Same system, same day. Only the question changed.",[88,89,90],"Denominator","What it answers","Honest?",[92,96,100,104],[93,94,95],"Products we attempted to match","How often the matcher produced something","No — excludes everything you never tried",[97,98,99],"Products in your catalogue","How much of your range you have competitor data for","Yes, and usually the one that matters",[101,102,103],"Products you actively compete on","Coverage where it has commercial consequence","Yes, and often the most useful",[105,106,107],"Products that exist on both sides","Matcher quality in isolation","Yes, for engineering. Not for a business claim.",{"type":109,"variant":110,"title":111,"text":112},"callout","warning","The 65% trap","A vendor reporting \"65% matched\" where the denominator is products the system attempted, rather than products in your catalogue, is quoting a number that can be improved by attempting fewer. We have been on both sides of this conversation. Always ask: 65% of what, counted how, and what happened to the rest?",{"type":80,"title":114,"paragraphs":115},"Two numbers, not one",[116,117,118,119],"Match rate alone cannot tell you whether a system is good, because it says nothing about whether the matches are right. Two measurements are needed and they trade against each other.","Precision: of the matches you published, what share are correct. This is the one that protects your credibility, and it is the one a pricing manager cares about, because a single visibly wrong comparison costs more trust than ten missing ones.","Recall: of the products that could have been matched, what share you found. This is coverage, and it is the one the person who bought the system asks about.","Pushing either to the limit destroys the other. A system that matches everything has terrible precision; a system that only matches barcodes has excellent precision and poor recall. The right balance comes from the cost asymmetry in lesson one, and for most commercial price monitoring that means favouring precision.",{"type":121,"title":122,"intro":123,"items":124},"steps","Building a sample you can trust","Two hundred hand-labelled pairs is enough to be useful and small enough that it actually gets done.",[125,128,131,134,137,140],{"title":126,"text":127},"Sample randomly, then stratify","Random across the whole catalogue, then top it up so each major category and each band is represented. Sampling only from matches you already made measures nothing.",{"title":129,"text":130},"Label by hand, with the pages open","Open both listings and decide. Allow three answers — same, different, genuinely ambiguous — and keep the ambiguous ones, because their share is itself a finding.",{"title":132,"text":133},"Have a second person label a slice","Fifty pairs, independently. Where two careful humans disagree is the ceiling on what any automated system can achieve, and it is usually lower than people expect.",{"title":135,"text":136},"Score the system against the labels","Precision and recall separately, and broken down by category. The aggregate hides everything interesting.",{"title":138,"text":139},"Hold the sample out","Do not tune thresholds against the set you measure with. If you do, you are measuring how well you memorised two hundred pairs.",{"title":141,"text":142},"Re-measure quarterly","Both catalogues drift. A weighting tuned a year ago is describing a world that has moved.",{"type":80,"title":144,"paragraphs":145},"Report it by category or you have reported nothing",[146,147],"An aggregate match rate averages a solved problem with an unsolved one and tells you about neither. Groceries and electronics carry barcodes and will pull the average up; furniture, fashion and bikes will pull it down, and those are often exactly the categories where someone is waiting on an answer.","A broken-down table — category, products, matched, precision — is immediately actionable. It shows where to spend effort, and it stops the conversation where someone quotes the aggregate at a category it does not describe.",{"type":84,"title":149,"intro":150,"headers":151,"rows":155},"Worked example: precision and recall from 200 labelled pairs","Two hundred pairs, labelled by a person, held back from everything used to tune the matcher. Of those, 120 are genuinely the same product and 80 are not, and the system proposed 110 matches. Both headline numbers come out of this one table and they move in opposite directions. Loosening the thresholds until recall reaches 95% would add about ten more true positives and take them from the borderline band, which is where the false ones live too — and for commercial pricing those 6 false positives already cost more than the 16 misses do.",[152,153,154],"","System said match","System said no match",[156,160,164],[157,158,159],"Genuinely the same product (120)","104 — true positives","16 — missed",[161,162,163],"Genuinely different (80)","6 — false, and expensive","74 — correctly rejected",[152,165,166],"Precision: 104 ÷ 110 = 94.5%","Recall: 104 ÷ 120 = 86.7%",{"type":168,"title":169,"intro":170,"items":171},"list","What usually goes wrong","Every item here produces a number that is higher than the truth and impossible to argue with.",[172,173,174,175,176,177],"Reporting one figure. Precision and recall trade against each other, so a single \"accuracy\" number can be improved by making the matcher either braver or more timid, and neither is information.","Labelling the sample from the system's own output. Pairs it never proposed cannot appear, so recall becomes unmeasurable by construction and comes out looking perfect.","Tuning on the evaluation set. The resulting number measures memorisation. Hold the 200 pairs back and do not look at them while adjusting anything.","Reporting one rate across all categories. Electronics with barcodes and fashion without them have completely different ceilings, and the blended figure describes neither of them.","Quietly changing the denominator when the number disappoints. Narrowing the scope explicitly and saying so is honest, and improving the input data is honest. The third option is not.","Not establishing the human ceiling. Have a second person label a slice: if two people disagree on 8% of pairs, a matcher scoring 92% is at the limit of the exercise rather than failing it.",{"type":80,"title":179,"paragraphs":180},"What to do with a number you do not like",[181,182,183,184],"Sometimes the honest measurement is bad. We have seen 18% on a bicycle catalogue — a genuinely hard category where the same frame is sold under different model names with different component specs, and the number was a correct description of the problem rather than a broken system.","There are three defensible responses and one that is not. You can narrow the scope, to the products where matching is reliable and the commercial stakes are real. You can change the input, by sourcing identifiers from the manufacturer rather than inferring them from retailer titles. You can accept it and report it, with the uncovered products clearly labelled.","What you cannot do is quietly change the denominator, which is the response the incentives push hardest towards. A match rate that improved because the question got easier is not an improvement, and the person it fools first is you.","The general rule, and it is the same one as everywhere else in this course: publish what you measured. A page, a report or a feed that says plainly where the data does not reach is worth more than one that pads the gap — because the first can be acted on, and the second will be believed.",[186,187,188,189],"A match rate without its denominator is not a claim. Ask: percent of what, and what happened to the rest?","Measure precision and recall separately — they trade against each other, and for commercial pricing, precision wins.","Two hundred hand-labelled pairs, held out from tuning, re-measured quarterly. Have a second person label a slice to find the human ceiling.","Report by category. When the number is bad, narrow the scope or improve the input — never quietly change the denominator.",{"text":191,"label":192,"url":193},"One question remains, and it is the one legal teams ask first. The final course covers what you are allowed to collect and how to answer for it.","Course 6: the legal and ethical side","/learn/web-scraping-legal-and-ethical",[195,201,206,212,217,223],{"slug":196,"navTitle":197,"title":198,"summary":199,"time":200,"needsAccount":73},"why-matching-is-the-hard-part","Why matching is hard","Why matching is the hard part","The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.","10 min",{"slug":202,"navTitle":203,"title":204,"summary":205,"time":72,"needsAccount":73},"identifiers-first-gtin-ean-mpn","Identifiers first","Identifiers first: GTIN, EAN, UPC and MPN","What each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.",{"slug":207,"navTitle":208,"title":209,"summary":210,"time":211,"needsAccount":73},"when-there-is-no-barcode","No barcode","Matching when there is no barcode","Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.","12 min",{"slug":213,"navTitle":214,"title":215,"summary":216,"time":211,"needsAccount":73},"variants-bundles-and-multipacks","Variants and packs","Variants, bundles and multipacks","The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.",{"slug":218,"navTitle":219,"title":220,"summary":221,"time":72,"needsAccount":222},"score-confidence-and-build-a-review-queue","Confidence and review","Confidence scores and a review queue worth using","Why one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.",true,{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":218,"navTitle":219,"title":220,"summary":221,"time":72,"needsAccount":222},null,1791047867364]