Match the same product across different sites· Lesson 4 of 6

Variants, bundles and multipacks

The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.

  • 12 min read
  • No account needed

Everything in the previous lesson can be done well and you will still produce confidently wrong matches, because the hardest cases are the ones where the titles genuinely almost agree. A six-pack and a single can. A drill with batteries and the same drill without. A 500ml bottle and a 750ml bottle of the identical product.

These are the matches that score highest and are most wrong, and they are the reason a price monitoring feed loses credibility in one afternoon.

The four shapes, and what each one requires

ShapeExampleComparable?What to do
Multipack6x330ml vs 1x330mlYes, per unitNormalise to a unit price and compare that
Size variant500ml vs 750ml of the same productPer unit, with careUnit price, but flag — bigger packs are priced differently on purpose
Non-price variantSame shoe, different colourYes, directlyTreat as one product unless the price actually differs by colour
BundleDrill + case + 2 batteries vs drill bodyNoDifferent products. Do not compare, and say why.

Normalise to a unit, and keep the raw value too

For packs and sizes the operation is mechanical: extract the pack count and the unit size from the title or the specifications, compute price per unit, and compare on that.

Two disciplines make this survive contact with reality. First, scrape the raw values and derive the unit price as a separate step, rather than trying to extract a computed value from the page. The page gives you price, currency and pack size; the division is yours. If you conflate them, a change in how the site formats its price becomes a change in your unit price with no visible cause.

Second, never discard the original. A pricing manager looking at a surprising row needs to see the shelf price and the pack size, not only the derived number. A unit price with no provenance is unauditable, and the first question anyone asks about a surprising comparison is "what was the actual price on the page".

Pack size extraction

Deliberately conservative. It returns nothing rather than guessing, because a wrong pack count is a wrong price.

python
import re

# 6x330ml, 6 x 330 ml, 24-pack, pack of 12
PACK = re.compile(
    r"(?:(\d+)\s*[x\u00d7]\s*(\d+(?:[.,]\d+)?)\s*(ml|l|g|kg|cl))"
    r"|(?:(\d+)[\s-]*pack)"
    r"|(?:pack\s+of\s+(\d+))",
    re.I,
)


def pack_size(title: str):
    """Return (count, unit_size, unit) or None. None means ask a human."""
    m = PACK.search(title)
    if not m:
        return None
    if m.group(1):
        return int(m.group(1)), float(m.group(2).replace(",", ".")), m.group(3).lower()
    count = m.group(4) or m.group(5)
    return int(count), None, None


def unit_price(price, title):
    pack = pack_size(title)
    if pack is None:
        return None          # unknown pack size is not the same as a pack of one
    return round(price / pack[0], 4)

The important line is the last return of None. Defaulting an unknown pack size to one is how a six-pack gets compared to a single can.

Variants are a modelling decision, not a matching one

Retailers disagree about what a product is. One publishes a single page with a size selector; another publishes eleven separate pages. Your matcher sees one row on one side and eleven on the other, and no amount of string comparison resolves that, because the disagreement is structural.

Decide the grain explicitly and hold it everywhere. Matching at the variant level is correct when price genuinely varies by variant — clothing sizes, phone storage capacity. Matching at the parent level is correct when it does not — colour variants of the same shoe at one price.

The failure mode to avoid is being inconsistent, where some sources are matched at parent level and others at variant level. That inflates your match count and silently double-counts a competitor in any index you compute. We have seen a run return a hundred rows that covered only forty-four distinct URLs, for exactly this reason, and the row count looked healthy the whole time.

Worked example: four listings that all look like a 50% undercut

Each row is a competitor listing scoring highly against the same product of yours — a 6-pack of 500 ml bottles at €11.94, which is €0.398 per 100 ml. Three of the four become comparable once the arithmetic is done. The fourth is not comparable at any price, and no amount of normalisation will make it so.

Their listingTheir pricePer 100 mlVerdict
6 × 500 ml€11.94€0.398identical — you are level
1 × 500 ml€2.49€0.498comparable — they are 25% dearer per unit, not 79% cheaper
12 × 500 ml€20.28€0.338comparable — genuinely 15% under you
Starter kit: 2 × 500 ml plus a dispenser€14.99not computableno match, with a reason — the dispenser has no price of its own

What usually goes wrong

Beyond the guards above, these are the decisions that let a pack-size problem reach a price.

  • Mixing grains across sources. If one competitor is tracked at variant level and another at parent level, the match count inflates and the same product is counted twice in every average you compute.
  • Storing the unit price and discarding the raw one. When a figure looks wrong, seeing both numbers side by side is the only way to separate a bad extraction from a bad conversion.
  • Letting the normalised figure move a price directly. A per-litre number is a comparison aid, not a shelf price; derive it, show it, and reprice from the comparison the rule actually intends.
  • Scraping a column with the same name as a derived one. A scraped unit_price blocks the rule meant to produce unit_price, and the rule then never runs — silently.
  • Treating "no match" as a failure to be engineered away. The starter-kit row is the correct output: unmatched with a reason is the only honest answer when one side includes something the other does not sell.
  • Falling back to headline price when the per-unit figure is unavailable. That second row reads as a 79% undercut and is in fact a 25% premium.

Guards worth having before any of this reaches a dashboard

  • Refuse to compare when either side's pack size is unknown. Unknown is not one.
  • Flag any match where the pack counts differ by more than a factor of four, even after normalisation
  • Detect bundle words — kit, set, bundle, with, plus, includes, body only, bare tool — and require them to agree on both sides
  • Record pack count and unit size as their own columns, so a bad extraction is visible rather than baked into a derived number
  • Carry a provenance note on every normalised comparison so a surprising row can be audited in one click

Next: what to do with everything that landed in the middle band.

Lesson 5: confidence scores and review queues

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.