[{"data":1,"prerenderedAt":1041},["ShallowReactive",2],{"learn-course-matching-products-across-sites":3,"learn-courses":830},{"slug":4,"order":5,"level":6,"time":7,"card_text":8,"seo":9,"hero":15,"outcomes":25,"who":34,"syllabus":45,"faq":48,"lessons":65},"matching-products-across-sites",5,"You have data from more than one site","6 lessons, about 70 minutes","Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.",{"title":10,"description":11,"keywords":12,"og_title":13,"og_description":14},"Matching Products Across Sites: A Free 6-Lesson Course","How to match the same product across different retailers. GTIN and EAN matching, fuzzy title matching, variants and multipacks, confidence scoring, review queues, and measuring match rate honestly. Free, ungated.","product matching, product data matching, gtin matching, ean barcode matching, fuzzy product matching, sku mapping, competitor price matching, product catalogue matching","A free course on matching the same product across different retailers","Six written lessons on product matching: identifiers first, fuzzy matching when there is no barcode, variants and multipacks, confidence scores and review queues.",{"badge":16,"title":17,"subtitle":18,"cta_primary":19,"cta_secondary":22},"Course five","Match the same product across different sites","A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.",{"label":20,"url":21},"Start with lesson one","/learn/matching-products-across-sites/why-matching-is-the-hard-part",{"label":23,"url":24},"See the data a run returns","/custom-scrapers",{"title":26,"items":27},"What you will be able to do",[28,29,30,31,32,33],"Explain why matching, not collection, is where price monitoring projects fail","Use GTIN, EAN, UPC and MPN correctly, including the checks that stop a bad identifier poisoning a match","Build a defensible fuzzy match for the majority of products that carry no usable barcode","Handle variants, bundles and multipacks without silently comparing a three-pack to a single unit","Attach a confidence score to every match and route the uncertain ones to a human","Measure your match rate with a denominator that tells the truth",{"title":35,"for_title":36,"for":37,"not_title":41,"not_for":42},"Who this is for","Written for",[38,39,40],"Pricing and e-commerce teams who have competitor data and cannot trust the comparisons","Developers building the matching layer and looking for the failure modes before they find them in production","Anyone evaluating a price monitoring vendor who wants to ask better questions than \"how many sites do you cover\"","Not written for",[43,44],"Teams who do not yet have data from more than one source — collection comes first, see course one or three","Anyone expecting a machine learning tutorial. The techniques here are mostly deterministic, because deterministic matches are the ones you can defend to a pricing manager.",{"title":46,"intro":47},"The six lessons","Lessons two and three are the mechanics. Lesson four is the one that catches most teams out, and lesson six is the one that stops you believing your own numbers.",{"badge":49,"title":50,"description":51,"items":52},"FAQ","Before you start","The questions that come up as soon as somebody looks closely at a match table.",[53,56,59,62],{"title":54,"description":55},"What match rate should I expect?","It depends almost entirely on the category, and anyone who quotes you a single number across all categories is not being careful. Electronics and groceries carry barcodes and match well. Fashion, furniture and bikes are hard because the same frame is sold under different model names with different component specs. We have seen a real project stall at 18% on bicycles, which was a correct measurement of a genuinely hard category rather than a broken system.",{"title":57,"description":58},"Can I just match on product title?","You can, and you will get a result that looks plausible and is wrong often enough to be dangerous. Lesson three is about doing it properly when there is no alternative, which there frequently is not. The important part is not the string algorithm — it is attaching a confidence score and refusing to auto-publish the weak ones.",{"title":60,"description":61},"Should a human be in the loop?","Yes, for the middle band. The strong matches do not need a person and the hopeless ones do not deserve one. The value of a review queue is entirely in how well you have sorted it, which is lesson five.",{"title":63,"description":64},"Is this specific to Scrapewise?","No. Lesson five mentions how we structure a review queue, and that lesson is labelled. The rest is method and applies to a spreadsheet, a Python script or any vendor's matching engine.",[66,196,327,458,575,713],{"slug":67,"nav_title":68,"title":69,"summary":70,"time":71,"needs_account":72,"seo":73,"blocks":77,"takeaways":187,"next_step":192},"why-matching-is-the-hard-part","Why matching is hard","Why matching is the hard part","The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.","10 min",false,{"title":74,"description":75,"keywords":76},"Why Product Matching Is Harder Than Scraping Prices","Retailers describe the same product differently on purpose. Why that makes matching the real bottleneck in price monitoring, and what a wrong match actually costs.","product matching, product data matching, price comparison accuracy, competitor price monitoring, catalogue matching",[78,83,88,121,126,132,142,173,182],{"type":79,"paragraphs":80},"prose",[81,82],"Teams start price monitoring projects worrying about collection. Can we read the pages, will we get blocked, how fresh will it be. Those are solvable and the answers are mostly engineering.","Then the data arrives and the real problem shows up: you have forty thousand rows from six retailers and no reliable way to say which of them refer to the same object. Every comparison, every price index, every alert depends on that mapping being right, and nothing in the collection layer helps you build it.",{"type":79,"title":84,"paragraphs":85},"The differences are deliberate",[86,87],"It is tempting to treat the variation in how retailers describe products as sloppiness. It is not. A retailer's product title is a marketing asset — it is tuned for their own search, their own SEO, their own customer. Two shops selling the identical item have active reasons to describe it differently, and a retailer with an exclusive-looking model name has a reason to make comparison harder.","So you are not cleaning up an accident. You are reconciling six catalogues that were each built for a different purpose, by people with no interest in them lining up.",{"type":89,"title":90,"intro":91,"headers":92,"rows":96},"table","One product, six retailers","A composite built from patterns we see constantly. Every column is a correct description of the same physical object.",[93,94,95],"Retailer","Title as published","What makes it hard",[97,101,105,109,113,117],[98,99,100],"A","Bosch Professional GSB 18V-55 Combi Drill (2x4.0Ah)","Battery spec in the title, pack contents implied",[102,103,104],"B","GSB18V-55 Cordless Hammer Drill Driver Kit","No brand word, model number unspaced, different category noun",[106,107,108],"C","Bosch Blue 18V Combi Drill + 2 Batteries + Case","Sub-brand instead of model, contents as a list",[110,111,112],"D","Bosch GSB 18V-55 (Body Only)","Looks like the same product. It is not — no batteries.",[114,115,116],"E","Akkuschrauber Bosch GSB 18V-55 Professional","Different language, different word order",[118,119,120],"F","Bosch 06019H5302","Manufacturer part number as the entire title",{"type":122,"variant":123,"title":124,"text":125},"callout","warning","Row D is the whole lesson","Of the six, five are the same purchasable thing and one is not. Retailer D is the body alone — no batteries, no charger, roughly half the price. A title-similarity match will score it extremely high against the others, and the result is an alert telling your pricing team that a competitor has undercut you by 48%. Somebody then reprices against a product that does not exist. The worst matches are not the ones that look obviously wrong.",{"type":79,"title":127,"paragraphs":128},"What a wrong match costs, specifically",[129,130,131],"A missed match costs you visibility. You do not see a competitor's price, your coverage is lower than you thought, and the index is computed over a smaller set. That is a real cost and it is recoverable.","A false match costs you money and credibility. It produces a confident, precise, wrong number, and that number goes into a repricing rule or a board slide. The first time a pricing manager finds one, they stop trusting the whole feed — and they are right to, because they have no way to know which other rows are wrong.","This asymmetry should drive the entire design. Being conservative and leaving a product unmatched is nearly always cheaper than guessing. Systems tuned to maximise match rate are tuned for the wrong thing.",{"type":133,"title":134,"intro":135,"items":136},"list","The five ways the same product diverges","Worth internalising, because each one needs a different handling strategy.",[137,138,139,140,141],"Naming — brand word present or absent, model number spaced or not, category noun varying","Language and locale — the same product in a different language, with different units and decimal separators","Pack and bundle — single unit, multipack, with-accessories, body-only","Variant — colour, size, capacity, which may or may not affect price","Identifier availability — some retailers publish a barcode, most do not, and some publish the wrong one",{"type":89,"title":143,"intro":144,"headers":145,"rows":149},"Worked example: what one false match costs","A single wrong pair, followed by an automatic repricing rule, over thirty days. Their listing is the 1-litre bottle and yours is the 2-litre; the two titles agree on brand, product and very nearly every word, which is why this match scored higher than most of your correct ones. You sell 40 units a day at €18.00 with €5.40 of margin, so the product costs you €12.60.",[146,147,148],"","Before the match","After the rule acts",[150,154,158,162,166,170],[151,152,153],"Their price, as matched","—","€9.95, being the 1-litre bottle",[155,156,157],"Your price","€18.00","€9.85",[159,160,161],"Margin per unit","€5.40","−€2.75",[163,164,165],"Units in thirty days","1,200","1,200, and probably more",[167,168,169],"Margin over thirty days","€6,480","−€3,300",[171,146,172],"Swing","€9,780, on one SKU",{"type":133,"title":174,"intro":175,"items":176},"What usually goes wrong","Matching goes wrong in the design rather than in the code, and nearly always in the same direction.",[177,178,179,180,181],"Tuning for match rate. It is the metric that rewards exactly the behaviour above, because the pairs it adds last are the ones most likely to be wrong.","Treating a missed match and a false match as the same size of error. One costs you visibility on a line; the other costs margin, quietly and continuously.","Assuming the collection layer helps. A flawless scrape of a wrong pair is a precisely wrong number, which is considerably harder to notice than a missing one.","Expecting retailers' titles to converge. They are marketing assets with a search term inside them, written differently on purpose, and no amount of cleaning turns them into a shared key.","Deferring the question until the data has arrived. Matching decides whether the feed is intelligence or fiction, so it belongs in the plan rather than in the clean-up.",{"type":79,"title":183,"paragraphs":184},"The order of operations for the rest of the course",[185,186],"Identifiers first, because when they exist they are nearly free and nearly certain. Then fuzzy matching for everything else, with a confidence score attached rather than a yes or no. Then variant and pack handling, which is the step that catches row D. Then a review queue for the band in the middle. Then measurement, so you know what you actually have rather than what you hope you have.","Skipping straight to fuzzy matching is the common mistake, and it means doing hard probabilistic work on products where a barcode would have answered the question outright.",[188,189,190,191],"Collection is engineering; matching is the bottleneck, and nothing in the collection layer helps with it.","Retailers describe products differently on purpose — their titles are marketing assets, not catalogue keys.","A missed match costs visibility. A false match costs money and credibility, and is often the highest-scoring one.","Design conservatively: unmatched is cheaper than wrongly matched, so do not tune for match rate.",{"text":193,"label":194,"url":195},"Next: the cheap certain wins — barcodes, part numbers, and the checks that keep them honest.","Lesson 2: identifiers first — GTIN, EAN, MPN","/learn/matching-products-across-sites/identifiers-first-gtin-ean-mpn",{"slug":197,"nav_title":198,"title":199,"summary":200,"time":201,"needs_account":72,"seo":202,"blocks":206,"takeaways":318,"next_step":323},"identifiers-first-gtin-ean-mpn","Identifiers first","Identifiers first: GTIN, EAN, UPC and MPN","What each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.","11 min",{"title":203,"description":204,"keywords":205},"GTIN, EAN, UPC and MPN: Product Identifiers for Matching","How product identifiers work, the check digit that validates a barcode, where retailers publish them, and the three ways a valid identifier still produces a wrong match.","gtin, ean barcode, upc code, mpn manufacturer part number, product identifier matching, barcode validation check digit",[207,211,246,250,255,262,268,275,304,313],{"type":79,"paragraphs":208},[209,210],"If two listings carry the same valid barcode, they are the same product, and you are done in one comparison. Nothing else in product matching is that cheap or that certain, so the first thing to build is the identifier path — even though it will only cover part of your catalogue.","The confusion worth clearing up first is that these are not competing standards. They are mostly the same number written at different lengths.",{"type":89,"title":212,"headers":213,"rows":218},"The identifiers you will meet",[214,215,216,217],"Identifier","What it is","Length","Reliability for matching",[219,224,228,232,237,242],[220,221,222,223],"GTIN","The umbrella term — GTIN-8, 12, 13 and 14 are all GTINs","8 to 14 digits","Highest, when valid",[225,226,227,223],"EAN","European article number, the common retail barcode","13 digits",[229,230,231,223],"UPC","North American equivalent; a GTIN-13 with a leading zero","12 digits",[233,234,235,236],"MPN","Manufacturer part number — the maker's own code","Free-form","High with the brand, meaningless without it",[238,239,240,241],"ASIN","Amazon's internal identifier","10 characters","Only within Amazon",[243,244,235,245],"Retailer SKU","The shop's own code","None across retailers",{"type":122,"variant":247,"title":248,"text":249},"note","UPC and EAN are the same number","A 12-digit UPC is a 13-digit EAN with a leading zero. If you store them in different columns and compare them as strings, every US-to-EU match will silently fail. Normalise everything to 13 or 14 digits on the way in and the problem disappears before it exists.",{"type":79,"title":251,"paragraphs":252},"Validate before you trust",[253,254],"Every GTIN carries a check digit, computed from the preceding digits. It exists so a scanner can tell a misread from a real code, and it works just as well for telling a scraped mistake from a real code.","Validating costs microseconds and removes a whole class of silent corruption: a truncated field, a number that was actually a model code, a cell that picked up a stray character. Any identifier that fails the check digit should be discarded rather than used, because a wrong identifier is more dangerous than no identifier — it will match confidently to something arbitrary.",{"type":256,"title":257,"intro":258,"language":259,"code":260,"caption":261},"code","Check digit validation","The standard GS1 modulo-10 calculation, normalised for any GTIN length.","python","def valid_gtin(code: str) -> bool:\n    \"\"\"True if code is a syntactically valid GTIN-8/12/13/14.\"\"\"\n    digits = \"\".join(c for c in code if c.isdigit())\n    if len(digits) not in (8, 12, 13, 14):\n        return False\n    body, check = digits[:-1], int(digits[-1])\n    # Weights alternate 3,1 from the rightmost body digit leftwards.\n    total = sum(int(d) * (3 if i % 2 == 0 else 1)\n                for i, d in enumerate(reversed(body)))\n    return (10 - total % 10) % 10 == check\n\n\ndef normalise(code: str) -> str | None:\n    \"\"\"Pad to 14 digits so UPC-12 and EAN-13 compare equal.\"\"\"\n    if not valid_gtin(code):\n        return None\n    return \"\".join(c for c in code if c.isdigit()).zfill(14)","Normalising to 14 digits is what makes a 12-digit UPC and the same product's 13-digit EAN compare as equal.",{"type":79,"title":263,"paragraphs":264},"Where retailers actually publish them",[265,266,267],"In descending order of how often it works: the JSON-LD Product block, where gtin13, gtin, sku and mpn are standard fields; a microdata itemprop; a specifications table on the page; and the product feed, if the retailer publishes one.","Expect patchy coverage and expect it to vary wildly by sector. Groceries and electronics are good. Fashion and furniture are poor, often because the item genuinely has no manufacturer barcode. That is not a gap you can engineer away, which is what lesson three exists for.","One practical note if you are pulling these out of structured data: an array in a JSON-LD path needs to be addressed as its own segment. A product with variants will put the barcodes under something like hasVariant with each entry carrying its own gtin13, and a path that does not account for the array returns the first one for everything.",{"type":133,"title":269,"intro":270,"items":271},"Three ways a valid identifier still produces a wrong match","The check digit proves the number is well-formed. It does not prove it is the right number for this listing.",[272,273,274],"Reused codes. Small manufacturers reuse barcodes across generations. The same EAN can be last year's model at one retailer and this year's at another.","Bundle codes. A retailer sells a drill plus a case under the bare drill's barcode because that is what their system had. The identifier is valid, the products are not equivalent.","Copy-paste errors in the retailer's own catalogue. A listing carries the barcode of a neighbouring variant. Rare individually, guaranteed across a hundred thousand rows.",{"type":89,"title":276,"intro":277,"headers":278,"rows":283},"Worked example: four codes, two products","Four identifiers exactly as four retailers publish them, covering two distinct products. All four are valid, all four are the same kind of number, and a plain string comparison matches none of them to anything at all. Left-padding everything to fourteen digits makes A equal B and C equal D — and that single step is what stops US and European catalogues failing to match in complete silence.",[279,280,281,282],"Source","Published as","Form","Normalised to GTIN-14",[284,289,292,297,301],[285,286,287,288],"A, a European retailer","4006381333931","EAN-13","04006381333931",[290,288,291,288],"B, another European retailer","GTIN-14, already padded",[293,294,295,296],"C, a US retailer","036000291452","UPC-A","00036000291452",[298,299,300,296],"D, the same US product elsewhere","0036000291452","the EAN-13 form of that UPC",[302,146,146,303],"A plain string comparison","matches none of the four to each other",{"type":133,"title":174,"intro":305,"items":306},"Identifiers are the one tier of evidence that can be certain, which is exactly why the failures here are so damaging.",[307,308,309,310,311,312],"Comparing identifiers as strings without padding. A 12-digit UPC and the 13-digit EAN of the same item differ by one leading zero and nothing else, and the match simply never happens.","Trusting an identifier because it is the right length. Validate the check digit: a transposed pair of digits produces a number that looks perfect and refers either to something else or to nothing.","Treating an MPN as globally unique. It is unique within a manufacturer, so it needs the brand beside it to mean anything, and two brands both using \"100\" is entirely ordinary.","Assuming coverage. Fashion and furniture frequently publish no barcode on the page at all, so an identifier-first strategy returns almost nothing there — a scoping fact rather than a bug.","Accepting an identifier match with no sanity check. Reused codes, bundle codes and plain catalogue errors all pass the check digit; comparing title overlap and the price ratio catches all three for almost no effort.","Taking the first identifier on the page. Related-product carousels and accessory blocks publish codes of their own, and the one you want is inside the main product block.",{"type":79,"title":314,"paragraphs":315},"The cheap guard that catches all three",[316,317],"After an identifier match, compare the two titles and the two prices. If the titles share almost no words, or the prices differ by more than roughly a factor of two, flag it rather than publishing it.","This is a handful of lines and it catches every one of the failure modes above, because all three produce a match that is confident and visibly implausible. An identifier match that disagrees with every other signal is not a strong match — it is a strong contradiction, and that is worth a human's attention rather than an auto-publish.",[319,320,321,322],"GTIN, EAN and UPC are the same number at different lengths — normalise to 14 digits or US-to-EU matches fail silently.","Validate the check digit and discard anything that fails. A wrong identifier is worse than none.","Coverage is patchy and sector-dependent; fashion and furniture often have no barcode at all.","Sanity-check every identifier match against title overlap and price ratio — reused codes, bundle codes and catalogue errors all pass the check digit.",{"text":324,"label":325,"url":326},"Next: the majority of your catalogue, which has no usable barcode at all.","Lesson 3: when there is no barcode","/learn/matching-products-across-sites/when-there-is-no-barcode",{"slug":328,"nav_title":329,"title":330,"summary":331,"time":332,"needs_account":72,"seo":333,"blocks":337,"takeaways":449,"next_step":454},"when-there-is-no-barcode","No barcode","Matching when there is no barcode","Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.","12 min",{"title":334,"description":335,"keywords":336},"Fuzzy Product Matching Without a Barcode","How to match products on title, brand and attributes when no identifier exists. Normalisation, blocking, multi-signal scoring and the mistakes that make fuzzy matching unreliable.","fuzzy product matching, string similarity matching, product title matching, brand normalisation, record linkage, entity resolution ecommerce",[338,342,358,361,367,398,403,431,440],{"type":79,"paragraphs":339},[340,341],"Most of your catalogue will not have a usable identifier on both sides. That is the normal case, not a data quality failure, and the work here is to build something that is defensibly right rather than something that feels clever.","The structure is four steps: normalise so the comparison is fair, block so you are not comparing everything to everything, score on several signals rather than one, and threshold into three bands instead of two.",{"type":343,"title":344,"items":345},"steps","The four steps",[346,349,352,355],{"title":347,"text":348},"Normalise both sides identically","Lowercase, strip punctuation, collapse whitespace, unify units (1l and 1000ml, 1kg and 1000g), remove retailer noise words like \"free delivery\" and \"new\". The single highest-value step is a brand alias table — Bosch Blue and Bosch Professional are the same brand, and no string algorithm will ever work that out.",{"title":350,"text":351},"Block before you compare","Comparing fifty thousand products against fifty thousand is 2.5 billion comparisons. Blocking cuts that by only comparing within a shared key — same brand, or same category, or sharing a rare token. You lose a small number of cross-block matches and gain a system that finishes.",{"title":353,"text":354},"Score on several signals, not one","Title similarity, brand agreement, numeric attribute agreement, pack size agreement and price ratio. Each contributes; none decides alone. The multi-signal part is what makes it defensible.",{"title":356,"text":357},"Threshold into three bands","Auto-accept, review, auto-reject. Two bands is the mistake — it forces every borderline case into a wrong answer.",{"type":122,"variant":247,"title":359,"text":360},"The algorithm matters less than the normalisation","Teams spend days choosing between Levenshtein, Jaro-Winkler, token-set ratio and embeddings. In practice the difference between them is small compared to the difference made by a brand alias table and consistent unit handling. Normalise properly with a mediocre algorithm and you will beat a sophisticated algorithm on raw strings.",{"type":79,"title":362,"paragraphs":363},"Numbers in titles are the strongest weak signal",[364,365,366],"Product titles are mostly marketing words, which are the least discriminating part of the string. The numbers are where the information is: 18V, 55Nm, 4.0Ah, 256GB, 1.5L.","A practical rule that outperforms its simplicity: extract every number-with-unit from both titles and treat disagreement as near-fatal. Two titles that are 90% similar as strings but say 128GB and 256GB are different products, full stop, and string similarity will never tell you that because one character differs out of forty.","The same logic applies in reverse. Two titles that share little text but agree on every numeric attribute, the brand, and sit within a few percent on price are very probably the same thing described by two different marketing departments.",{"type":89,"title":368,"intro":369,"headers":370,"rows":374},"A workable signal weighting","A starting point, not a law. Tune on a hand-labelled sample — see lesson six.",[371,372,373],"Signal","Weight","Notes",[375,379,382,386,390,394],[376,377,378],"Brand agreement","Gate","Different brands should not match at all. Treat as a filter, not a score.",[380,377,381],"Numeric attribute agreement","Any disagreement on a spec number is near-fatal",[383,384,385],"Title token overlap","40%","After normalisation, ignoring stopwords",[387,388,389],"Model or part number match","35%","Huge when present, absent on most listings",[391,392,393],"Pack size agreement","15%","Covered properly in lesson four",[395,396,397],"Price plausibility","10%","Weak signal, real one. A 4x gap on a supposed match is a warning.",{"type":79,"title":399,"paragraphs":400},"Why price is a signal and also a trap",[401,402],"Price belongs in the score because an implausible ratio is genuine evidence against a match. It gets a small weight for a reason, though: you are building this system to detect price differences, so weighting price heavily means the system becomes least confident exactly when a competitor does something interesting.","Use it as a tiebreaker and a warning flag, never as a primary. The rule we settled on is that a price ratio beyond about three to one demotes a match to review regardless of how well everything else scored — because at that ratio, body-only versus full-kit is a more likely explanation than a genuine undercut.",{"type":89,"title":404,"intro":405,"headers":406,"rows":410},"Worked example: why blocking is not optional","Your 1,200 products against one competitor's 29,000 listings. Comparing every pair is nearly thirty-five million comparisons for one competitor on one run, and the great majority of them are a kettle against a garden hose. Each divisor below is just the number of distinct values in that field, so the arithmetic carries over to your own catalogue with your own figures.",[407,408,409],"Approach","Candidate pairs","What it removes",[411,415,419,423,427],[412,413,414],"Every product against every listing","1,200 × 29,000 = 34,800,000","nothing",[416,417,418],"Block on brand, with 60 brands in play","÷ 60 = about 580,000","every pair whose brands disagree",[420,421,422],"Block on brand and category, with 6 categories","÷ 6 = about 96,000","the kettle against the garden hose",[424,425,426],"Also require a shared number in the title, roughly 1 pair in 13","÷ 13 = about 7,400","most of the remaining near-misses",[428,429,430],"Of those, how many a person ever sees","a few hundred","which is the whole reason for three bands rather than one threshold",{"type":133,"title":174,"intro":432,"items":433},"Separate from the scoring mistakes above, these are the ones that come from how the matcher is run and evaluated.",[434,435,436,437,438,439],"Skipping the blocking step and reaching for a faster machine. Thirty-five million comparisons is a design problem, not a performance problem, and the index that fixes it is an afternoon's work.","Blocking on a field that is not always populated. A block on category silently discards every listing where the category was not extracted, and those pairs never get the chance to match at all.","Evaluating on the pairs the system found. The pairs it missed are invisible by construction, so recall can only be measured against a hand-labelled sample.","Running the matcher over every row every night. Confirmed pairs are facts; re-deriving them nightly means the answer can change without anyone having changed anything.","Having no band for \"do not know\". Forcing every pair to match or not match converts an honest uncertainty into a wrong number on a dashboard.","Reading the score as a probability. It is a weighted sum whose weights you chose: 0.87 is an ordering, not a seven-in-eight chance of being right.",{"type":133,"title":441,"items":442},"Mistakes that make fuzzy matching untrustworthy",[443,444,445,446,447,448],"One threshold instead of three bands, which forces a wrong answer on every borderline pair","No brand gate, so two different manufacturers' similar products match on generic words","Treating a 95% string score as certainty when the 5% is the capacity or the pack size","Normalising one side and not the other, which quietly halves every score","Tuning the thresholds against the same examples you built them from, which measures memorisation rather than accuracy","Never re-checking. A weighting tuned on last year's catalogue drifts as both catalogues change.",[450,451,452,453],"Normalise, block, score on multiple signals, threshold into three bands.","A brand alias table and consistent unit handling beat a better string algorithm.","Numbers in titles carry the information. Treat spec-number disagreement as near-fatal.","Keep price as a low-weight tiebreaker — weighting it highly makes you least confident exactly when something interesting happens.",{"text":455,"label":456,"url":457},"Next: the failure that survives every technique above — packs, bundles and variants.","Lesson 4: variants, bundles and multipacks","/learn/matching-products-across-sites/variants-bundles-and-multipacks",{"slug":459,"nav_title":460,"title":461,"summary":462,"time":332,"needs_account":72,"seo":463,"blocks":467,"takeaways":566,"next_step":571},"variants-bundles-and-multipacks","Variants and packs","Variants, bundles and multipacks","The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.",{"title":464,"description":465,"keywords":466},"Matching Product Variants, Bundles and Multipacks","Why pack size and variant differences produce the most dangerous false matches in price comparison, how to normalise to a unit price, and when to refuse to compare.","product variant matching, multipack price comparison, unit price normalisation, bundle pricing comparison, pack size data",[468,472,500,503,509,514,520,549,558],{"type":79,"paragraphs":469},[470,471],"Everything in the previous lesson can be done well and you will still produce confidently wrong matches, because the hardest cases are the ones where the titles genuinely almost agree. A six-pack and a single can. A drill with batteries and the same drill without. A 500ml bottle and a 750ml bottle of the identical product.","These are the matches that score highest and are most wrong, and they are the reason a price monitoring feed loses credibility in one afternoon.",{"type":89,"title":473,"headers":474,"rows":479},"The four shapes, and what each one requires",[475,476,477,478],"Shape","Example","Comparable?","What to do",[480,485,490,495],[481,482,483,484],"Multipack","6x330ml vs 1x330ml","Yes, per unit","Normalise to a unit price and compare that",[486,487,488,489],"Size variant","500ml vs 750ml of the same product","Per unit, with care","Unit price, but flag — bigger packs are priced differently on purpose",[491,492,493,494],"Non-price variant","Same shoe, different colour","Yes, directly","Treat as one product unless the price actually differs by colour",[496,497,498,499],"Bundle","Drill + case + 2 batteries vs drill body","No","Different products. Do not compare, and say why.",{"type":122,"variant":123,"title":501,"text":502},"The bundle row is not a normalisation problem","The first three rows can be solved with arithmetic. The fourth cannot. A kit and a bare body are different purchasable things and no unit conversion makes them comparable. The correct output is \"no match\", with a note. Teams that try to solve bundles by discounting the accessories are inventing a number, and an invented number in a price feed is worse than a gap in one.",{"type":79,"title":504,"paragraphs":505},"Normalise to a unit, and keep the raw value too",[506,507,508],"For packs and sizes the operation is mechanical: extract the pack count and the unit size from the title or the specifications, compute price per unit, and compare on that.","Two disciplines make this survive contact with reality. First, scrape the raw values and derive the unit price as a separate step, rather than trying to extract a computed value from the page. The page gives you price, currency and pack size; the division is yours. If you conflate them, a change in how the site formats its price becomes a change in your unit price with no visible cause.","Second, never discard the original. A pricing manager looking at a surprising row needs to see the shelf price and the pack size, not only the derived number. A unit price with no provenance is unauditable, and the first question anyone asks about a surprising comparison is \"what was the actual price on the page\".",{"type":256,"title":510,"intro":511,"language":259,"code":512,"caption":513},"Pack size extraction","Deliberately conservative. It returns nothing rather than guessing, because a wrong pack count is a wrong price.","import re\n\n# 6x330ml, 6 x 330 ml, 24-pack, pack of 12\nPACK = re.compile(\n    r\"(?:(\\d+)\\s*[x\\u00d7]\\s*(\\d+(?:[.,]\\d+)?)\\s*(ml|l|g|kg|cl))\"\n    r\"|(?:(\\d+)[\\s-]*pack)\"\n    r\"|(?:pack\\s+of\\s+(\\d+))\",\n    re.I,\n)\n\n\ndef pack_size(title: str):\n    \"\"\"Return (count, unit_size, unit) or None. None means ask a human.\"\"\"\n    m = PACK.search(title)\n    if not m:\n        return None\n    if m.group(1):\n        return int(m.group(1)), float(m.group(2).replace(\",\", \".\")), m.group(3).lower()\n    count = m.group(4) or m.group(5)\n    return int(count), None, None\n\n\ndef unit_price(price, title):\n    pack = pack_size(title)\n    if pack is None:\n        return None          # unknown pack size is not the same as a pack of one\n    return round(price / pack[0], 4)","The important line is the last return of None. Defaulting an unknown pack size to one is how a six-pack gets compared to a single can.",{"type":79,"title":515,"paragraphs":516},"Variants are a modelling decision, not a matching one",[517,518,519],"Retailers disagree about what a product is. One publishes a single page with a size selector; another publishes eleven separate pages. Your matcher sees one row on one side and eleven on the other, and no amount of string comparison resolves that, because the disagreement is structural.","Decide the grain explicitly and hold it everywhere. Matching at the variant level is correct when price genuinely varies by variant — clothing sizes, phone storage capacity. Matching at the parent level is correct when it does not — colour variants of the same shoe at one price.","The failure mode to avoid is being inconsistent, where some sources are matched at parent level and others at variant level. That inflates your match count and silently double-counts a competitor in any index you compute. We have seen a run return a hundred rows that covered only forty-four distinct URLs, for exactly this reason, and the row count looked healthy the whole time.",{"type":89,"title":521,"intro":522,"headers":523,"rows":528},"Worked example: four listings that all look like a 50% undercut","Each row is a competitor listing scoring highly against the same product of yours — a 6-pack of 500 ml bottles at €11.94, which is €0.398 per 100 ml. Three of the four become comparable once the arithmetic is done. The fourth is not comparable at any price, and no amount of normalisation will make it so.",[524,525,526,527],"Their listing","Their price","Per 100 ml","Verdict",[529,534,539,544],[530,531,532,533],"6 × 500 ml","€11.94","€0.398","identical — you are level",[535,536,537,538],"1 × 500 ml","€2.49","€0.498","comparable — they are 25% dearer per unit, not 79% cheaper",[540,541,542,543],"12 × 500 ml","€20.28","€0.338","comparable — genuinely 15% under you",[545,546,547,548],"Starter kit: 2 × 500 ml plus a dispenser","€14.99","not computable","no match, with a reason — the dispenser has no price of its own",{"type":133,"title":174,"intro":550,"items":551},"Beyond the guards above, these are the decisions that let a pack-size problem reach a price.",[552,553,554,555,556,557],"Mixing grains across sources. If one competitor is tracked at variant level and another at parent level, the match count inflates and the same product is counted twice in every average you compute.","Storing the unit price and discarding the raw one. When a figure looks wrong, seeing both numbers side by side is the only way to separate a bad extraction from a bad conversion.","Letting the normalised figure move a price directly. A per-litre number is a comparison aid, not a shelf price; derive it, show it, and reprice from the comparison the rule actually intends.","Scraping a column with the same name as a derived one. A scraped unit_price blocks the rule meant to produce unit_price, and the rule then never runs — silently.","Treating \"no match\" as a failure to be engineered away. The starter-kit row is the correct output: unmatched with a reason is the only honest answer when one side includes something the other does not sell.","Falling back to headline price when the per-unit figure is unavailable. That second row reads as a 79% undercut and is in fact a 25% premium.",{"type":133,"title":559,"items":560},"Guards worth having before any of this reaches a dashboard",[561,562,563,564,565],"Refuse to compare when either side's pack size is unknown. Unknown is not one.","Flag any match where the pack counts differ by more than a factor of four, even after normalisation","Detect bundle words — kit, set, bundle, with, plus, includes, body only, bare tool — and require them to agree on both sides","Record pack count and unit size as their own columns, so a bad extraction is visible rather than baked into a derived number","Carry a provenance note on every normalised comparison so a surprising row can be audited in one click",[567,568,569,570],"The highest-scoring false matches are pack, variant and bundle cases, because the titles genuinely almost agree.","Packs and sizes normalise with arithmetic. Bundles do not — the honest output is no match plus a reason.","Scrape raw price, currency and pack size; derive the unit price separately and keep the original visible.","Pick a grain — variant level or parent level — and hold it across every source, or your match count inflates and double-counts.",{"text":572,"label":573,"url":574},"Next: what to do with everything that landed in the middle band.","Lesson 5: confidence scores and review queues","/learn/matching-products-across-sites/score-confidence-and-build-a-review-queue",{"slug":576,"nav_title":577,"title":578,"summary":579,"time":201,"needs_account":580,"seo":581,"blocks":585,"takeaways":704,"next_step":709},"score-confidence-and-build-a-review-queue","Confidence and review","Confidence scores and a review queue worth using","Why one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.",true,{"title":582,"description":583,"keywords":584},"Product Match Confidence Scores and Review Queues","How to score match confidence without conflating certainty about the price with certainty about the match, where to set auto-accept and reject thresholds, and how to order a human review queue by value.","match confidence score, human in the loop matching, product match review queue, data labelling workflow, match threshold tuning",[586,589,595,619,622,628,638,645,689,698],{"type":79,"paragraphs":587},[588],"A matching system that outputs yes or no is throwing away the most useful thing it knows, which is how sure it is. Once the score survives into the output, you can route by it: publish the top band, discard the bottom, and give a person the middle, which is usually a small fraction of the catalogue and the only part where human time pays for itself.",{"type":79,"title":590,"paragraphs":591},"One score is not enough",[592,593,594],"This is a mistake we shipped and had to unpick. A single confidence number tends to conflate two separate questions: how certain are we that we read this price correctly, and how certain are we that these two products are the same thing.","They are independent. A price can be read perfectly from a page whose product is only probably the right match. A price can be a guess from a page that is unambiguously the right product. Collapsing those into one number produces outputs that are genuinely confusing to the person reading them — a high score on a row that is obviously wrong, and no way to tell which half was the problem.","Carry two: extraction confidence and match confidence. They route differently. Low extraction confidence is an engineering ticket. Low match confidence is a review item. Sending either to the wrong queue wastes the time of whoever receives it.",{"type":89,"title":596,"intro":597,"headers":598,"rows":603},"Three bands, by match confidence","Starting values. Lesson six is how to replace them with measured ones.",[599,600,601,602],"Band","Score","Action","Typical share",[604,609,614],[605,606,607,608],"Auto-accept","Valid identifier, or score above ~0.9 with all gates passed","Publish","Varies hugely by category",[610,611,612,613],"Review","~0.6 to ~0.9, or any gate conflict","Human queue, ordered by value","Aim for under 10% of the catalogue",[615,616,617,618],"Reject","Below ~0.6, or a brand or spec gate failed","Discard, log the near-misses","The rest",{"type":122,"variant":123,"title":620,"text":621},"Set the thresholds from the cost of each error, not from the shape of the histogram","The question is not where the scores cluster. It is what a false match costs you compared to a missed one. If a wrong comparison can trigger an automated reprice, the accept threshold needs to be high enough that you would defend every row in it to a finance director. If the output is an analyst's weekly reading, you can afford to be looser.",{"type":79,"title":623,"paragraphs":624},"Order the queue by value, not by score",[625,626,627],"This is the difference between a review queue that works and one that quietly stops being opened.","Sorting by confidence puts your reviewer on the most ambiguous rows in the catalogue, which are also frequently the least important ones — obscure products nobody competes on. They spend an hour on hard calls with no business consequence, and the queue feels like a punishment.","Sort by potential impact instead: your sales volume for that product, multiplied by the price gap the match implies. Now the first twenty rows are the ones where being wrong actually costs something. An hour on that list is worth having, and the reviewer can see why.",{"type":133,"title":629,"intro":630,"items":631},"What a review row needs on screen","The goal is a decision in under ten seconds. Anything that forces a tab switch breaks the rhythm.",[632,633,634,635,636,637],"Both titles, with the differing tokens highlighted","Both images, side by side — the fastest signal a human has","Both prices, plus the implied ratio","Both pack sizes and the key spec numbers, aligned","Why the matcher was unsure, named explicitly: \"pack size disagreement\", \"no shared model number\"","Three buttons: same, not the same, cannot tell. The third one is essential and routinely omitted.",{"type":122,"variant":639,"title":640,"text":641,"cta":642},"product","Where this lives in Scrapewise","Match review runs against the collected rows with both sides' titles, prices and images in one view, and decisions feed back as overrides that survive the next run. The thresholds and the queue ordering are yours to set — the part we do not do for you is deciding what a false match is worth in your business, because that number is not ours to guess.",{"label":643,"url":644},"Create a free account","https://portal.scrapewise.ai/register",{"type":89,"title":646,"intro":647,"headers":648,"rows":655},"Worked example: the same queue, ordered two ways","Five pairs waiting, one hour of somebody's attention, two orderings. Score ordering puts the reviewer on the pairs the system is least sure about; value ordering — monthly units times the price gap — puts them on the pairs where being wrong costs something. The two disagree completely: the score ordering spends the hour on €26.80 of exposure and the value ordering spends it on €7,860. Rows A and B never get looked at under the second ordering, and that is the correct outcome.",[649,600,650,651,652,653,654],"Pair","Monthly units","Price gap","At stake","Rank by score","Rank by value",[656,664,671,678,683],[657,658,659,660,661,662,663],"A. Niche accessory","0.52","4","€1.20","€4.80","1","5",[665,666,667,668,669,670,659],"B. Discontinued line","0.58","11","€2.00","€22.00","2",[672,673,674,675,676,677,670],"C. Mid-range appliance","0.71","90","€14.00","€1,260","3",[679,680,681,669,682,659,662],"D. Top-selling TV","0.74","300","€6,600",[684,685,686,687,688,663,677],"E. Spare part","0.80","25","€3.50","€87.50",{"type":133,"title":174,"intro":690,"items":691},"A review queue is a product with exactly one user, and it fails for the reasons products fail.",[692,693,694,695,696,697],"Carrying one score. Extraction confidence and match confidence answer different questions, and a blended number cannot tell you whether to re-scrape or to review.","Setting thresholds from the shape of the histogram. The gap in the distribution is an artefact of your own weighting; the thresholds belong where the cost of a false match overtakes the cost of a missed one.","Ordering by score. The hour goes on the cheapest pairs in the catalogue, exactly as in the table.","A review row without the evidence on it. If the reviewer has to open two tabs to decide, throughput collapses and the queue stops being opened at all.","Not persisting the decision. A queue that presents the same pair again next week is a queue that gets abandoned in week three.","Keying the decision on a URL. Retailers change URLs, so key it on product identity — otherwise every confirmed match quietly expires the next time somebody tidies a slug.",{"type":79,"title":699,"paragraphs":700},"Decisions are an asset — persist them",[701,702,703],"Every human judgement is expensive and permanently useful. Store it as an override keyed on the two product identities, not on row ids that change between runs, and make sure the next run reads it before it scores anything.","The failure mode here is brutal and common: a review queue that re-presents the same hundred pairs every week because the decisions were never written back. People stop opening it within a month, entirely reasonably.","Two refinements worth adding once the basics work. Record who decided and when, so a disputed comparison can be traced. And re-surface old decisions when the underlying titles change materially — a product that was correctly matched last year may have been replaced by a new generation under the same listing.",[705,706,707,708],"Carry two scores — extraction confidence and match confidence. Conflating them produces outputs nobody can act on.","Three bands, with thresholds set from the cost of each error rather than the shape of the distribution.","Order the review queue by volume times price gap, not by score, or your reviewer spends the hour on products nobody competes on.","Persist decisions keyed on product identity. A queue that repeats itself stops being opened.",{"text":710,"label":711,"url":712},"Finally: measuring what you have built, with a denominator that does not flatter you.","Lesson 6: measure your match rate honestly","/learn/matching-products-across-sites/measure-your-match-rate-honestly",{"slug":714,"nav_title":715,"title":716,"summary":717,"time":201,"needs_account":72,"seo":718,"blocks":722,"takeaways":821,"next_step":826},"measure-your-match-rate-honestly","Measure it honestly","Measure your match rate honestly","The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.",{"title":719,"description":720,"keywords":721},"How to Measure Product Match Rate Correctly","Why match rate is usually quoted against a flattering denominator, how to build a hand-labelled sample, measuring precision and recall separately, and reporting a low number well.","match rate measurement, precision and recall matching, product matching accuracy, price monitoring coverage, data quality metrics",[723,726,750,753,760,782,787,805,814],{"type":79,"paragraphs":724},[725],"\"We match 94% of products\" is a sentence with no meaning attached until somebody says what the denominator was. It is also the single most common way both vendors and internal teams mislead themselves — usually without intending to, because the flattering denominator is the one that is easiest to compute.",{"type":89,"title":727,"intro":728,"headers":729,"rows":733},"Four denominators, four very different numbers","Same system, same day. Only the question changed.",[730,731,732],"Denominator","What it answers","Honest?",[734,738,742,746],[735,736,737],"Products we attempted to match","How often the matcher produced something","No — excludes everything you never tried",[739,740,741],"Products in your catalogue","How much of your range you have competitor data for","Yes, and usually the one that matters",[743,744,745],"Products you actively compete on","Coverage where it has commercial consequence","Yes, and often the most useful",[747,748,749],"Products that exist on both sides","Matcher quality in isolation","Yes, for engineering. Not for a business claim.",{"type":122,"variant":123,"title":751,"text":752},"The 65% trap","A vendor reporting \"65% matched\" where the denominator is products the system attempted, rather than products in your catalogue, is quoting a number that can be improved by attempting fewer. We have been on both sides of this conversation. Always ask: 65% of what, counted how, and what happened to the rest?",{"type":79,"title":754,"paragraphs":755},"Two numbers, not one",[756,757,758,759],"Match rate alone cannot tell you whether a system is good, because it says nothing about whether the matches are right. Two measurements are needed and they trade against each other.","Precision: of the matches you published, what share are correct. This is the one that protects your credibility, and it is the one a pricing manager cares about, because a single visibly wrong comparison costs more trust than ten missing ones.","Recall: of the products that could have been matched, what share you found. This is coverage, and it is the one the person who bought the system asks about.","Pushing either to the limit destroys the other. A system that matches everything has terrible precision; a system that only matches barcodes has excellent precision and poor recall. The right balance comes from the cost asymmetry in lesson one, and for most commercial price monitoring that means favouring precision.",{"type":343,"title":761,"intro":762,"items":763},"Building a sample you can trust","Two hundred hand-labelled pairs is enough to be useful and small enough that it actually gets done.",[764,767,770,773,776,779],{"title":765,"text":766},"Sample randomly, then stratify","Random across the whole catalogue, then top it up so each major category and each band is represented. Sampling only from matches you already made measures nothing.",{"title":768,"text":769},"Label by hand, with the pages open","Open both listings and decide. Allow three answers — same, different, genuinely ambiguous — and keep the ambiguous ones, because their share is itself a finding.",{"title":771,"text":772},"Have a second person label a slice","Fifty pairs, independently. Where two careful humans disagree is the ceiling on what any automated system can achieve, and it is usually lower than people expect.",{"title":774,"text":775},"Score the system against the labels","Precision and recall separately, and broken down by category. The aggregate hides everything interesting.",{"title":777,"text":778},"Hold the sample out","Do not tune thresholds against the set you measure with. If you do, you are measuring how well you memorised two hundred pairs.",{"title":780,"text":781},"Re-measure quarterly","Both catalogues drift. A weighting tuned a year ago is describing a world that has moved.",{"type":79,"title":783,"paragraphs":784},"Report it by category or you have reported nothing",[785,786],"An aggregate match rate averages a solved problem with an unsolved one and tells you about neither. Groceries and electronics carry barcodes and will pull the average up; furniture, fashion and bikes will pull it down, and those are often exactly the categories where someone is waiting on an answer.","A broken-down table — category, products, matched, precision — is immediately actionable. It shows where to spend effort, and it stops the conversation where someone quotes the aggregate at a category it does not describe.",{"type":89,"title":788,"intro":789,"headers":790,"rows":793},"Worked example: precision and recall from 200 labelled pairs","Two hundred pairs, labelled by a person, held back from everything used to tune the matcher. Of those, 120 are genuinely the same product and 80 are not, and the system proposed 110 matches. Both headline numbers come out of this one table and they move in opposite directions. Loosening the thresholds until recall reaches 95% would add about ten more true positives and take them from the borderline band, which is where the false ones live too — and for commercial pricing those 6 false positives already cost more than the 16 misses do.",[146,791,792],"System said match","System said no match",[794,798,802],[795,796,797],"Genuinely the same product (120)","104 — true positives","16 — missed",[799,800,801],"Genuinely different (80)","6 — false, and expensive","74 — correctly rejected",[146,803,804],"Precision: 104 ÷ 110 = 94.5%","Recall: 104 ÷ 120 = 86.7%",{"type":133,"title":174,"intro":806,"items":807},"Every item here produces a number that is higher than the truth and impossible to argue with.",[808,809,810,811,812,813],"Reporting one figure. Precision and recall trade against each other, so a single \"accuracy\" number can be improved by making the matcher either braver or more timid, and neither is information.","Labelling the sample from the system's own output. Pairs it never proposed cannot appear, so recall becomes unmeasurable by construction and comes out looking perfect.","Tuning on the evaluation set. The resulting number measures memorisation. Hold the 200 pairs back and do not look at them while adjusting anything.","Reporting one rate across all categories. Electronics with barcodes and fashion without them have completely different ceilings, and the blended figure describes neither of them.","Quietly changing the denominator when the number disappoints. Narrowing the scope explicitly and saying so is honest, and improving the input data is honest. The third option is not.","Not establishing the human ceiling. Have a second person label a slice: if two people disagree on 8% of pairs, a matcher scoring 92% is at the limit of the exercise rather than failing it.",{"type":79,"title":815,"paragraphs":816},"What to do with a number you do not like",[817,818,819,820],"Sometimes the honest measurement is bad. We have seen 18% on a bicycle catalogue — a genuinely hard category where the same frame is sold under different model names with different component specs, and the number was a correct description of the problem rather than a broken system.","There are three defensible responses and one that is not. You can narrow the scope, to the products where matching is reliable and the commercial stakes are real. You can change the input, by sourcing identifiers from the manufacturer rather than inferring them from retailer titles. You can accept it and report it, with the uncovered products clearly labelled.","What you cannot do is quietly change the denominator, which is the response the incentives push hardest towards. A match rate that improved because the question got easier is not an improvement, and the person it fools first is you.","The general rule, and it is the same one as everywhere else in this course: publish what you measured. A page, a report or a feed that says plainly where the data does not reach is worth more than one that pads the gap — because the first can be acted on, and the second will be believed.",[822,823,824,825],"A match rate without its denominator is not a claim. Ask: percent of what, and what happened to the rest?","Measure precision and recall separately — they trade against each other, and for commercial pricing, precision wins.","Two hundred hand-labelled pairs, held out from tuning, re-measured quarterly. Have a second person label a slice to find the human ceiling.","Report by category. When the number is bad, narrow the scope or improve the input — never quietly change the denominator.",{"text":827,"label":828,"url":829},"One question remains, and it is the one legal teams ask first. The final course covers what you are allowed to collect and how to answer for it.","Course 6: the legal and ethical side","/learn/web-scraping-legal-and-ethical",[831,883,923,961,1000,1008],{"order":832,"slug":833,"title":834,"subtitle":835,"cardText":836,"level":837,"time":838,"lessonCount":839,"lessons":840},1,"competitor-price-monitoring","Build a competitor price monitoring pipeline","Price monitoring looks like a scraping problem for about a week. Then you discover that scraping was the easy part, and the project actually lives or dies on which competitors you picked, whether their listings are really the same product as yours, and whether anyone notices the morning the feed comes back half empty. This course is those eight decisions, in the order you have to make them.","From \"we check three competitors by hand on Mondays\" to a feed you trust enough to reprice from. The eight decisions in order, including the two that quietly ruin most projects.","No coding required","8 lessons, about 90 minutes",8,[841,847,852,857,862,868,873,878],{"slug":842,"navTitle":843,"title":844,"summary":845,"time":846,"needsAccount":72},"what-is-competitor-price-monitoring","What it actually is","What competitor price monitoring actually is","The four stages of a price pipeline, why only two of them are scraping, and the one question to ask before you build anything.","9 min",{"slug":848,"navTitle":849,"title":850,"summary":851,"time":201,"needsAccount":72},"choose-competitors-and-skus","Choosing what to track","Choosing which competitors and which SKUs to track","How to build a list that is small enough to afford and large enough to matter, using margin at risk rather than gut feel.",{"slug":853,"navTitle":854,"title":855,"summary":856,"time":332,"needsAccount":580},"find-competitor-product-urls","Finding product URLs","Finding every competitor product URL without copying them by hand","Four ways to get a competitor's full product URL list, ranked by how much work they are, and what to do when none of them work.",{"slug":858,"navTitle":859,"title":860,"summary":861,"time":332,"needsAccount":580},"extract-price-stock-and-shipping","Extracting the fields","Getting price, stock and shipping off the page","Which fields to extract, why the sale price is two fields and not one, and the four ways a price appears on a page.",{"slug":863,"navTitle":864,"title":865,"summary":866,"time":867,"needsAccount":72},"match-listings-to-your-catalogue","Matching to your catalogue","Matching competitor listings to your own catalogue","The stage that decides whether your feed is intelligence or fiction, and the denominator trick that makes bad match rates look good.","13 min",{"slug":869,"navTitle":870,"title":871,"summary":872,"time":201,"needsAccount":72},"schedule-runs-and-catch-silent-failure","Scheduling and data quality","Scheduling runs and catching silent data loss","How often to actually check, and the four alerts that catch a degrading feed before someone reprices from it.",{"slug":874,"navTitle":875,"title":876,"summary":877,"time":846,"needsAccount":580},"export-to-sheets-bi-and-erp","Getting the data out","Getting the data into Sheets, BI or your ERP","Four delivery routes ranked by how likely they are to actually get used, and the column contract that stops downstream jobs breaking.",{"slug":879,"navTitle":880,"title":881,"summary":882,"time":332,"needsAccount":72},"turn-price-data-into-repricing-rules","From data to decisions","Turning price data into repricing decisions","Why \"match the cheapest\" destroys margin, what a rule needs besides a competitor price, and how to start without automating anything.",{"order":884,"slug":885,"title":886,"subtitle":887,"cardText":888,"level":889,"time":890,"lessonCount":891,"lessons":892},2,"ai-agent-web-data-mcp","Give your AI agent live web data via MCP","Ask an assistant what a product costs today and you will usually get a number. It is often wrong, and it is always wrong in the same way: the model is reconstructing a plausible price from training data rather than looking at a page. This course is about closing that gap properly — what the Model Context Protocol actually is, how to wire a server into a client, how to design tools a model can use without hand-holding, and what to put in place before an agent spends your money.","Your agent is confidently wrong about prices because it has never seen one. What MCP is, how to connect a server, how to design tools a model can actually use, and the guardrails you need before you let it loose.","Comfortable editing a config file","6 lessons, about 60 minutes",6,[893,898,903,908,913,918],{"slug":894,"navTitle":895,"title":896,"summary":897,"time":846,"needsAccount":72},"what-is-mcp","What MCP is","What MCP actually is, in plain terms","The Model Context Protocol described without jargon: what problem it solves, its three primitives, and when it is the wrong tool.",{"slug":899,"navTitle":900,"title":901,"summary":902,"time":71,"needsAccount":72},"why-agents-get-live-data-wrong","Why agents get it wrong","Why your agent's answer about a price is wrong","Four distinct failure modes that all look identical from the outside, and how to tell which one you have before you try to fix it.",{"slug":904,"navTitle":905,"title":906,"summary":907,"time":71,"needsAccount":72},"connect-an-mcp-server","Connecting a server","Connecting an MCP server and proving it works","The config for local and remote servers, the four things that go wrong, and how to verify the tools registered rather than assuming.",{"slug":909,"navTitle":910,"title":911,"summary":912,"time":201,"needsAccount":72},"design-tools-an-agent-can-use","Designing usable tools","Designing tools an agent can actually use","A connected server is not a useful server. The model only sees your tool names, descriptions and parameter schemas, so those three things are the entire user interface. Here is what makes a tool get called correctly and what makes it get ignored.",{"slug":914,"navTitle":915,"title":916,"summary":917,"time":332,"needsAccount":580},"give-an-agent-a-scraper","Giving an agent a scraper","Giving an agent a real price feed","A worked example. Connect the ScrapeWise MCP server to a client, let the agent read a live scraper's output, and watch where the hand-off between \"the data is right\" and \"the answer is right\" actually breaks.",{"slug":919,"navTitle":920,"title":921,"summary":922,"time":201,"needsAccount":72},"guardrails-cost-and-untrusted-content","Guardrails and cost","Guardrails, cost control and untrusted content","Live web access turns an agent into something that can spend money and read text written by strangers. Neither is a reason not to do it. Both are reasons to put limits in before you need them.",{"order":924,"slug":925,"title":926,"subtitle":927,"cardText":928,"level":929,"time":7,"lessonCount":891,"lessons":930},3,"product-data-api","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.","Comfortable with HTTP and JSON",[931,936,941,946,951,956],{"slug":932,"navTitle":933,"title":934,"summary":935,"time":71,"needsAccount":72},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.",{"slug":937,"navTitle":938,"title":939,"summary":940,"time":71,"needsAccount":72},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":942,"navTitle":943,"title":944,"summary":945,"time":332,"needsAccount":580},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.",{"slug":947,"navTitle":948,"title":949,"summary":950,"time":201,"needsAccount":72},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.",{"slug":952,"navTitle":953,"title":954,"summary":955,"time":201,"needsAccount":72},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"slug":957,"navTitle":958,"title":959,"summary":960,"time":332,"needsAccount":580},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",{"order":962,"slug":963,"title":964,"subtitle":965,"cardText":966,"level":967,"time":968,"lessonCount":891,"lessons":969},4,"keep-scrapers-alive","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.","You already have something running","6 lessons, about 65 minutes",[970,975,980,985,990,995],{"slug":971,"navTitle":972,"title":973,"summary":974,"time":71,"needsAccount":72},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.",{"slug":976,"navTitle":977,"title":978,"summary":979,"time":332,"needsAccount":72},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.",{"slug":981,"navTitle":982,"title":983,"summary":984,"time":201,"needsAccount":72},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.",{"slug":986,"navTitle":987,"title":988,"summary":989,"time":201,"needsAccount":72},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":991,"navTitle":992,"title":993,"summary":994,"time":201,"needsAccount":72},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":996,"navTitle":997,"title":998,"summary":999,"time":71,"needsAccount":72},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"order":5,"slug":4,"title":17,"subtitle":18,"cardText":8,"level":6,"time":7,"lessonCount":891,"lessons":1001},[1002,1003,1004,1005,1006,1007],{"slug":67,"navTitle":68,"title":69,"summary":70,"time":71,"needsAccount":72},{"slug":197,"navTitle":198,"title":199,"summary":200,"time":201,"needsAccount":72},{"slug":328,"navTitle":329,"title":330,"summary":331,"time":332,"needsAccount":72},{"slug":459,"navTitle":460,"title":461,"summary":462,"time":332,"needsAccount":72},{"slug":576,"navTitle":577,"title":578,"summary":579,"time":201,"needsAccount":580},{"slug":714,"navTitle":715,"title":716,"summary":717,"time":201,"needsAccount":72},{"order":891,"slug":1009,"title":1010,"subtitle":1011,"cardText":1012,"level":1013,"time":1014,"lessonCount":5,"lessons":1015},"web-scraping-legal-and-ethical","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.","No legal background assumed","5 lessons, about 55 minutes",[1016,1021,1026,1031,1036],{"slug":1017,"navTitle":1018,"title":1019,"summary":1020,"time":71,"needsAccount":72},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.",{"slug":1022,"navTitle":1023,"title":1024,"summary":1025,"time":201,"needsAccount":72},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.",{"slug":1027,"navTitle":1028,"title":1029,"summary":1030,"time":201,"needsAccount":72},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"slug":1032,"navTitle":1033,"title":1034,"summary":1035,"time":201,"needsAccount":72},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.",{"slug":1037,"navTitle":1038,"title":1039,"summary":1040,"time":332,"needsAccount":72},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.",1791047866726]