[{"data":1,"prerenderedAt":245},["ShallowReactive",2],{"learn-lesson-matching-products-across-sites-score-confidence-and-build-a-review-queue":3},{"course":4,"lesson":67,"index":212,"outline":213,"prev":243,"next":244},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"matching-products-across-sites",5,"You have data from more than one site","6 lessons, about 70 minutes","Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Matching Products Across Sites: A Free 6-Lesson Course","How to match the same product across different retailers. GTIN and EAN matching, fuzzy title matching, variants and multipacks, confidence scoring, review queues, and measuring match rate honestly. Free, ungated.","product matching, product data matching, gtin matching, ean barcode matching, fuzzy product matching, sku mapping, competitor price matching, product catalogue matching","A free course on matching the same product across different retailers","Six written lessons on product matching: identifiers first, fuzzy matching when there is no barcode, variants and multipacks, confidence scores and review queues.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course five","Match the same product across different sites","A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.",{"label":21,"url":22},"Start with lesson one","/learn/matching-products-across-sites/why-matching-is-the-hard-part",{"label":24,"url":25},"See the data a run returns","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Explain why matching, not collection, is where price monitoring projects fail","Use GTIN, EAN, UPC and MPN correctly, including the checks that stop a bad identifier poisoning a match","Build a defensible fuzzy match for the majority of products that carry no usable barcode","Handle variants, bundles and multipacks without silently comparing a three-pack to a single unit","Attach a confidence score to every match and route the uncertain ones to a human","Measure your match rate with a denominator that tells the truth",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Pricing and e-commerce teams who have competitor data and cannot trust the comparisons","Developers building the matching layer and looking for the failure modes before they find them in production","Anyone evaluating a price monitoring vendor who wants to ask better questions than \"how many sites do you cover\"","Not written for",[44,45],"Teams who do not yet have data from more than one source — collection comes first, see course one or three","Anyone expecting a machine learning tutorial. The techniques here are mostly deterministic, because deterministic matches are the ones you can defend to a pricing manager.",{"title":47,"intro":48},"The six lessons","Lessons two and three are the mechanics. Lesson four is the one that catches most teams out, and lesson six is the one that stops you believing your own numbers.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions that come up as soon as somebody looks closely at a match table.",[54,57,60,63],{"title":55,"description":56},"What match rate should I expect?","It depends almost entirely on the category, and anyone who quotes you a single number across all categories is not being careful. Electronics and groceries carry barcodes and match well. Fashion, furniture and bikes are hard because the same frame is sold under different model names with different component specs. We have seen a real project stall at 18% on bicycles, which was a correct measurement of a genuinely hard category rather than a broken system.",{"title":58,"description":59},"Can I just match on product title?","You can, and you will get a result that looks plausible and is wrong often enough to be dangerous. Lesson three is about doing it properly when there is no alternative, which there frequently is not. The important part is not the string algorithm — it is attaching a confidence score and refusing to auto-publish the weak ones.",{"title":61,"description":62},"Should a human be in the loop?","Yes, for the middle band. The strong matches do not need a person and the hopeless ones do not deserve one. The value of a review queue is entirely in how well you have sorted it, which is lesson five.",{"title":64,"description":65},"Is this specific to Scrapewise?","No. Lesson five mentions how we structure a review queue, and that lesson is labelled. The rest is method and applies to a spreadsheet, a Python script or any vendor's matching engine.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":203,"next_step":208},"score-confidence-and-build-a-review-queue","Confidence and review","Confidence scores and a review queue worth using","Why one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.","11 min",true,{"title":75,"description":76,"keywords":77},"Product Match Confidence Scores and Review Queues","How to score match confidence without conflating certainty about the price with certainty about the match, where to set auto-accept and reject thresholds, and how to order a human review queue by value.","match confidence score, human in the loop matching, product match review queue, data labelling workflow, match threshold tuning",[79,83,89,114,119,125,136,143,187,197],{"type":80,"paragraphs":81},"prose",[82],"A matching system that outputs yes or no is throwing away the most useful thing it knows, which is how sure it is. Once the score survives into the output, you can route by it: publish the top band, discard the bottom, and give a person the middle, which is usually a small fraction of the catalogue and the only part where human time pays for itself.",{"type":80,"title":84,"paragraphs":85},"One score is not enough",[86,87,88],"This is a mistake we shipped and had to unpick. A single confidence number tends to conflate two separate questions: how certain are we that we read this price correctly, and how certain are we that these two products are the same thing.","They are independent. A price can be read perfectly from a page whose product is only probably the right match. A price can be a guess from a page that is unambiguously the right product. Collapsing those into one number produces outputs that are genuinely confusing to the person reading them — a high score on a row that is obviously wrong, and no way to tell which half was the problem.","Carry two: extraction confidence and match confidence. They route differently. Low extraction confidence is an engineering ticket. Low match confidence is a review item. Sending either to the wrong queue wastes the time of whoever receives it.",{"type":90,"title":91,"intro":92,"headers":93,"rows":98},"table","Three bands, by match confidence","Starting values. Lesson six is how to replace them with measured ones.",[94,95,96,97],"Band","Score","Action","Typical share",[99,104,109],[100,101,102,103],"Auto-accept","Valid identifier, or score above ~0.9 with all gates passed","Publish","Varies hugely by category",[105,106,107,108],"Review","~0.6 to ~0.9, or any gate conflict","Human queue, ordered by value","Aim for under 10% of the catalogue",[110,111,112,113],"Reject","Below ~0.6, or a brand or spec gate failed","Discard, log the near-misses","The rest",{"type":115,"variant":116,"title":117,"text":118},"callout","warning","Set the thresholds from the cost of each error, not from the shape of the histogram","The question is not where the scores cluster. It is what a false match costs you compared to a missed one. If a wrong comparison can trigger an automated reprice, the accept threshold needs to be high enough that you would defend every row in it to a finance director. If the output is an analyst's weekly reading, you can afford to be looser.",{"type":80,"title":120,"paragraphs":121},"Order the queue by value, not by score",[122,123,124],"This is the difference between a review queue that works and one that quietly stops being opened.","Sorting by confidence puts your reviewer on the most ambiguous rows in the catalogue, which are also frequently the least important ones — obscure products nobody competes on. They spend an hour on hard calls with no business consequence, and the queue feels like a punishment.","Sort by potential impact instead: your sales volume for that product, multiplied by the price gap the match implies. Now the first twenty rows are the ones where being wrong actually costs something. An hour on that list is worth having, and the reviewer can see why.",{"type":126,"title":127,"intro":128,"items":129},"list","What a review row needs on screen","The goal is a decision in under ten seconds. Anything that forces a tab switch breaks the rhythm.",[130,131,132,133,134,135],"Both titles, with the differing tokens highlighted","Both images, side by side — the fastest signal a human has","Both prices, plus the implied ratio","Both pack sizes and the key spec numbers, aligned","Why the matcher was unsure, named explicitly: \"pack size disagreement\", \"no shared model number\"","Three buttons: same, not the same, cannot tell. The third one is essential and routinely omitted.",{"type":115,"variant":137,"title":138,"text":139,"cta":140},"product","Where this lives in Scrapewise","Match review runs against the collected rows with both sides' titles, prices and images in one view, and decisions feed back as overrides that survive the next run. The thresholds and the queue ordering are yours to set — the part we do not do for you is deciding what a false match is worth in your business, because that number is not ours to guess.",{"label":141,"url":142},"Create a free account","https://portal.scrapewise.ai/register",{"type":90,"title":144,"intro":145,"headers":146,"rows":153},"Worked example: the same queue, ordered two ways","Five pairs waiting, one hour of somebody's attention, two orderings. Score ordering puts the reviewer on the pairs the system is least sure about; value ordering — monthly units times the price gap — puts them on the pairs where being wrong costs something. The two disagree completely: the score ordering spends the hour on €26.80 of exposure and the value ordering spends it on €7,860. Rows A and B never get looked at under the second ordering, and that is the correct outcome.",[147,95,148,149,150,151,152],"Pair","Monthly units","Price gap","At stake","Rank by score","Rank by value",[154,162,169,176,181],[155,156,157,158,159,160,161],"A. Niche accessory","0.52","4","€1.20","€4.80","1","5",[163,164,165,166,167,168,157],"B. Discontinued line","0.58","11","€2.00","€22.00","2",[170,171,172,173,174,175,168],"C. Mid-range appliance","0.71","90","€14.00","€1,260","3",[177,178,179,167,180,157,160],"D. Top-selling TV","0.74","300","€6,600",[182,183,184,185,186,161,175],"E. Spare part","0.80","25","€3.50","€87.50",{"type":126,"title":188,"intro":189,"items":190},"What usually goes wrong","A review queue is a product with exactly one user, and it fails for the reasons products fail.",[191,192,193,194,195,196],"Carrying one score. Extraction confidence and match confidence answer different questions, and a blended number cannot tell you whether to re-scrape or to review.","Setting thresholds from the shape of the histogram. The gap in the distribution is an artefact of your own weighting; the thresholds belong where the cost of a false match overtakes the cost of a missed one.","Ordering by score. The hour goes on the cheapest pairs in the catalogue, exactly as in the table.","A review row without the evidence on it. If the reviewer has to open two tabs to decide, throughput collapses and the queue stops being opened at all.","Not persisting the decision. A queue that presents the same pair again next week is a queue that gets abandoned in week three.","Keying the decision on a URL. Retailers change URLs, so key it on product identity — otherwise every confirmed match quietly expires the next time somebody tidies a slug.",{"type":80,"title":198,"paragraphs":199},"Decisions are an asset — persist them",[200,201,202],"Every human judgement is expensive and permanently useful. Store it as an override keyed on the two product identities, not on row ids that change between runs, and make sure the next run reads it before it scores anything.","The failure mode here is brutal and common: a review queue that re-presents the same hundred pairs every week because the decisions were never written back. People stop opening it within a month, entirely reasonably.","Two refinements worth adding once the basics work. Record who decided and when, so a disputed comparison can be traced. And re-surface old decisions when the underlying titles change materially — a product that was correctly matched last year may have been replaced by a new generation under the same listing.",[204,205,206,207],"Carry two scores — extraction confidence and match confidence. Conflating them produces outputs nobody can act on.","Three bands, with thresholds set from the cost of each error rather than the shape of the distribution.","Order the review queue by volume times price gap, not by score, or your reviewer spends the hour on products nobody competes on.","Persist decisions keyed on product identity. A queue that repeats itself stops being opened.",{"text":209,"label":210,"url":211},"Finally: measuring what you have built, with a denominator that does not flatter you.","Lesson 6: measure your match rate honestly","/learn/matching-products-across-sites/measure-your-match-rate-honestly",4,[214,221,226,232,237,238],{"slug":215,"navTitle":216,"title":217,"summary":218,"time":219,"needsAccount":220},"why-matching-is-the-hard-part","Why matching is hard","Why matching is the hard part","The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.","10 min",false,{"slug":222,"navTitle":223,"title":224,"summary":225,"time":72,"needsAccount":220},"identifiers-first-gtin-ean-mpn","Identifiers first","Identifiers first: GTIN, EAN, UPC and MPN","What each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.",{"slug":227,"navTitle":228,"title":229,"summary":230,"time":231,"needsAccount":220},"when-there-is-no-barcode","No barcode","Matching when there is no barcode","Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.","12 min",{"slug":233,"navTitle":234,"title":235,"summary":236,"time":231,"needsAccount":220},"variants-bundles-and-multipacks","Variants and packs","Variants, bundles and multipacks","The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":239,"navTitle":240,"title":241,"summary":242,"time":72,"needsAccount":220},"measure-your-match-rate-honestly","Measure it honestly","Measure your match rate honestly","The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.",{"slug":233,"navTitle":234,"title":235,"summary":236,"time":231,"needsAccount":220},{"slug":239,"navTitle":240,"title":241,"summary":242,"time":72,"needsAccount":220},1791047867355]