USE CASE

Product Matching Software for Messy Catalogs

The same product is listed under a different title, a different SKU and a different spec sheet on every site that sells it. ScrapeWise matches them, then delivers the matched rows to your pricing team — no matching model to build, train or maintain.

PAIN POINTS

Why Product Matching Breaks

  1. 01

    No Shared Identifier

    Most retailers do not publish a GTIN or EAN, and the ones that do often publish the wrong one. Matching on identifiers alone quietly misses most of the catalog.

  2. 02

    Titles Disagree by Design

    Every retailer writes its own title. Pack size, colour name, model year and bundled accessories all move around, so exact string matching finds almost nothing.

  3. 03

    Variants Collapse Into One

    A 500 g bag and a 2 kg bag share the title, the image and often the SKU stem. Match them together and your price comparison is wrong, not just incomplete.

  4. 04

    Confidence Is Not Reported

    Tools that return a match without a score hide their own uncertainty. You cannot review what you cannot see, so bad matches reach the pricing decision unchallenged.

  5. 05

    The Model Needs Owners

    Building matching in-house means someone owns the training data, the thresholds and the drift — permanently, not as a one-off project.

EUR 0.15

Per 1,000 delivered pages — matched and validated, no plan and no compute units

36

Ready-made scraper endpoints you can call directly for marketplace and retailer data

5

Free requests on every new account — test the match quality before you commit

HOW IT WORKS

How Product Matching Works Here

01

Collect the Competing Listings

We scrape the retailers and marketplaces you care about, capturing titles, attributes, images, identifiers, price and availability from each product page.

02

Match Against Your Catalog

Listings are normalised and matched to your SKUs on several signals at once — text, structured attributes, images and identifiers — rather than a single brittle rule.

03

Deliver Matched, Validated Rows

Matched pairs land in the portal and in your systems by REST API or CSV, on your schedule, with uncertain matches flagged for a human to confirm.

One Product, Four Listings, Four Different Answers

This is the actual sequence a single SKU goes through — WHEY-VAN-500G, a 500 g vanilla whey protein, against four competitor listings that all look like it.

  1. Normalise Before Comparing

    "Whey Protein Isolat Vanille 500g Beutel", "Vanilla Whey 0,5 kg Pouch" and "Whey Protein Vanilla 2 kg Dose" are parsed into structured attributes: brand, flavour, net weight in grams, container type, language. 0,5 kg and 500 g become the same number. 2 kg does not.

  2. Score Across Every Available Signal

    Each listing is scored on the signals it actually publishes. The amazon.de listing carries a GTIN, an image and a full spec block, so it scores 0.98. The idealo.de listing publishes a title and nothing else, so it scores 0.56 — not because it is a bad match, but because there is nothing to confirm it with.

  3. Decide, Hold or Escalate

    Two listings clear the threshold and are delivered as matches. The 2 kg pack is held as a separate variant rather than folded in. The title-only listing goes to review, where a person confirms or rejects it — before it can move a price.

The Signal Stack, and What Each Signal Misses Alone

Every matching approach that fails in production failed because it trusted one signal. These are the four we combine, and the specific blind spot each one has on its own.

Identifiers — GTIN, EAN, MPN, ASIN

The strongest signal when present and correct. The problem is coverage: most retailers publish no identifier at all, and a meaningful share of the ones that do publish the manufacturer's code for a different pack size or a superseded model. Identifier-only matching is high precision and low recall — it is right when it fires, and it rarely fires.

  • gtin
  • ean
  • mpn
  • asin
  • retailer_sku

Title Text

The only signal that is always there, and the least trustworthy. Retailers pad titles with bundled accessories, drop the pack size, translate the colour name, or keep last year's model number. Fuzzy string similarity alone will happily match a 500 g bag to a 2 kg tub because the strings are 90% identical.

  • title
  • brand
  • normalised_title

Structured Attributes

Spec tables, JSON-LD blocks and breadcrumb categories give you net weight, colour, size, material and model year as actual fields. This is what separates variants. Its blind spot is availability — thin retailer pages publish nothing structured, which is exactly where the title-only cases come from.

  • net_weight
  • colour
  • size
  • model_year
  • category

Images

Image similarity catches products whose titles share no useful words — a rebranded listing, a translated title, a marketplace seller writing their own description. It cannot separate variants: the 500 g bag and the 2 kg tub are routinely photographed with the identical manufacturer press shot.

  • image_url
  • image_hash

Confidence Is Reported, Not Hidden

A matcher that returns a match with no score is asserting certainty it does not have. Ours reports the score, and the borderline cases stop in a queue instead of flowing straight into a price change.

  • High Confidence — Delivered

    Multiple independent signals agree, including at least one that can separate variants. These rows go straight into your pricing system. In the worked example above, amazon.de at 0.98 and kaufland.de at 0.94.

  • Borderline — Review Queue

    The evidence points one way but nothing confirms the variant. The pair stays visible in the portal with both listings side by side, so a person confirms or rejects it in seconds. idealo.de at 0.56 — a title and nothing else.

  • Variant Conflict — Held

    The listing is clearly the same product line but a different pack size, colour or model year. It is not a match and it is not a rejection, so it is held as a separate variant rather than silently folded in. otto.de, 2 kg against your 500 g.

The Variant Trap

Folding variants together does not make your coverage look better. It makes your price comparison confidently wrong, which is worse than having no comparison at all.

  • Pack Size

    A 2 kg tub at EUR 49.90 next to your 500 g bag at EUR 16.90 reads as a competitor undercutting you by 66% per kilo. Fold them together and your repricing rule chases a price that does not exist.

  • Colour and Finish

    Colourways share the title, the spec table and often the SKU stem. Some sell at a premium, some at clearance. Matched as one product, the clearance colour drags your entire price position down.

  • Model Year and Bundle

    Last year's model and this year's, or the body-only listing and the one with the accessory kit, are different products at different prices. Treated as one, you are benchmarking against a discontinued SKU.

What You Send, What You Receive

You supply your catalog and the sites you want matched against it. You receive rows that are already linked to your SKUs.

You give
  • Your catalogrequired

    SKU, title, brand and whatever attributes you already hold. CSV upload or API push.

    WHEY-VAN-500G, Whey Protein Vanilla 500 g, BrandCo
  • Sites to match againstrequired

    Retailers and marketplaces by domain. Category pages, search pages or a URL list.

    amazon.dekaufland.deidealo.deotto.de
  • Identifiers, if you have themoptional

    GTIN, EAN or MPN improve precision where they exist. Matching does not depend on them.

    4012345678901
  • Variant policyoptional
    • Keep variants separate

    On by default. Pack sizes, colourways and model years are never folded into one match.

  • Refresh scheduleoptional
    • Daily
    • Weekly
    • On demand

    Matches are re-checked on your cadence, so retailer title changes do not quietly rot yesterday's matches.

You get

11 columns each

  • your_sku
  • competitor_domain
  • competitor_url
  • matched_title
  • match_confidence
  • match_signals
  • variant_status
  • price
  • currency
  • availability
  • checked_at

Delivered by REST API, CSV or scheduled export. Every row carries its confidence score and the signals that produced it.

The Matched Row You Actually Receive

One row per confirmed competitor listing, keyed to your SKU. This is the WHEY-VAN-500G example as it lands in your system.

your_skucompetitor_urlmatched_titleconfidencevariant_statuspriceavailability
WHEY-VAN-500Gamazon.de/dp/...Whey Protein Isolat Vanille 500g Beutel0.98exactEUR 18.49in stock
WHEY-VAN-500Gkaufland.de/product/...Vanilla Whey 0,5 kg Pouch0.94exactEUR 17.95in stock
WHEY-VAN-500Gotto.de/p/...Whey Protein Vanilla 2 kg Dose0.71variant_heldEUR 49.90in stock
WHEY-VAN-500Gidealo.de/offer/...Vanilla Whey Protein0.56in_reviewEUR 16.20unknown
BUILD VS BUY

Build the Matcher, or Buy the Matches

Feature
Building It In-House
Scrapewise
Training data
You label thousands of pairs, then relabel as the catalog changes
None to own — you review flagged pairs, not a dataset
Threshold policy
Someone picks a cutoff and defends it every time a match is wrong
Confidence is reported per row; borderline pairs stop for review
Data collection
Scrapers, proxies and anti-bot handling are a separate project
Included — 36 ready-made endpoints, we own the maintenance
Retailer title changes
Silent drift; you find out through a mispriced SKU
Re-matched on your schedule; drift is our problem
Variant handling
The failure mode you discover last and fix hardest
Variants kept separate by default
Cost you compare
Engineer-weeks, proxy gigabytes, compute units
EUR 0.15 per 1,000 delivered pages, 5 free requests to test
BENEFITS

What You Get Instead

Matched Rows, Not Raw Pages

Matched Rows, Not Raw Pages

You receive the competitor product already linked to yours, with price and availability attached. No parsing step, no reconciliation step, no spreadsheet in the middle.

Every Match Is Reviewable

Every Match Is Reviewable

Matches carry a confidence signal and stay visible in the portal, so your team can confirm the uncertain ones instead of discovering them through a mispriced SKU.

Nothing to Maintain

Nothing to Maintain

We own the scrapers, the proxies, the anti-bot handling and the matching itself. When a retailer changes its layout or its titles, that is our problem, not your sprint.

Stop Rebuilding the Matcher

If your pricing team is comparing the wrong products, the pricing decision is wrong before anyone looks at it. Get matched competitor rows delivered instead — no plan, no compute units, just a balance you top up.

FAQ

Product Matching FAQs

How product matching works at ScrapeWise, what it costs, and when you should build it yourself instead.

Product matching software identifies when two listings on different websites are the same physical product, even though the title, SKU, spec sheet and images all differ. It is the step that makes competitor price comparison possible: without it you are comparing your product against something that merely looks similar. Good matching combines several signals — text, structured attributes, images and identifiers such as GTIN or EAN — because any one of them on its own misses too much of the catalog.