Build a competitor price monitoring pipeline· Lesson 4 of 8

Getting price, stock and shipping off the page

Which fields to extract, why the sale price is two fields and not one, and the four ways a price appears on a page.

  • 12 min read

Now the scraping. This is the part everyone pictures when they imagine price monitoring, and it is genuinely the least interesting stage — which is good news, because it means you can treat it as plumbing and spend your attention elsewhere.

Two decisions matter here. What to extract, and how to make it survive the competitor redesigning their site.

Price is never one number

The most common beginner mistake is a single column called price. A product page typically shows a crossed-out original, a current selling price, sometimes a member price, sometimes a per-unit price underneath, and sometimes a basket-only price that is not displayed at all until checkout. Collapsing that into one number throws away the thing you most want to know, which is whether you are being undercut by a permanent reposition or by a three-day promotion.

Extract them separately and derive the comparison later. A discount percentage calculated at extraction time is a value you cannot recheck; two prices stored side by side can be re-derived any way you like, forever.

The field list

The first five are non-negotiable. The rest earn their place case by case.

FieldTake it?Note
priceAlwaysThe current selling price as displayed, digits only, no symbol
list_priceAlwaysThe crossed-out original where one is shown; empty is a valid answer
currencyAlwaysStop guessing from the domain. Multi-market stores serve several currencies on one host.
availabilityAlwaysThe exact words on the page, not your interpretation of them
urlAlwaysThe canonical URL, so you can click through and verify a surprising row
pack_size / unitUsuallyThe single most valuable matching field in grocery, chemicals and anything sold by volume
shipping_costSometimesOften only visible at checkout. If it is on the page, take it; do not build a basket flow to get it.
seller_nameMarketplaces onlyMeaningless on a single-brand store; essential on one where third parties list
ean / mpnIf shownWorth more than every other field combined when it comes to matching
review_countRarelyInteresting, never actionable, and it changes every day so it bloats your change history

Four places a price hides

Knowing which of these a site uses tells you immediately how fragile your extraction will be.

In the HTML, in a predictable element. The classic case. A CSS selector reads it, and it breaks the day the competitor changes their theme.

In JSON-LD structured data. Many stores publish a machine-readable Product block in the page source, for Google. When it is there and it is accurate, it is the best source available: it is explicitly a contract with search engines, so it changes far less often than the visual markup. It is also frequently incomplete, wrong on sale prices, or absent on exactly the pages you care about — so verify, do not assume.

In an internal API call the page makes after loading. The HTML arrives with a price-shaped hole in it and JavaScript fills it in. Any extraction that reads raw HTML gets nothing; you need something that runs the page's JavaScript first.

In an image. Rare, deliberate, and a signal that the site does not want to be read. Treat it as a competitor you track manually.

Selectors versus a model that reads the page

A CSS selector is precise, free, fast and brittle. It encodes the competitor's current HTML structure into your pipeline, and it breaks on redesign — not with an error, which would be helpful, but usually by returning nothing or by returning the wrong element that happens to sit where the old one did.

The alternative is to describe the field in words and have a model find it on the rendered page. That survives a redesign, because a price is still visibly a price after the CSS changes. It costs more per page and it is not infallible — a model can also pick the wrong element, and it will do so with complete confidence.

In practice the deciding factor is maintenance headcount. If nobody owns fixing broken selectors within a day, selectors are a false economy: the cheapest extraction in the world is worthless on the mornings it returns nothing.

Worked example: one product page, eight columns

A single listing, as a shopper sees it on the left and as it should land in your table on the right. Every judgement has been deferred: nothing here is interpreted, only recorded. Note what is absent — there is no discount column, because 1 − 199 ÷ 249 is 20.1% today and will still be 20.1% whenever anyone asks, recomputed from two columns that were written down.

What the page showsColumnStored value
€249.00, struck throughlist_price249.00
€199.00, in redsale_price199.00
The € symbol before the figurecurrencyEUR
"In stock — 3 left"availability_textIn stock — 3 left
nothing on the pageavailability_classleft empty, derived downstream
"Free delivery over €50"shipping_textFree delivery over €50
read at 06:14 on 14 Marchcaptured_at2026-03-14T06:14:00Z
the page the figures came fromsource_urlhttps://…/p/12345

What usually goes wrong

Extraction failures are rarely loud. Four of these five produce a table that looks entirely reasonable and is wrong.

  • One price column. The page showed two numbers and the table kept one, so nobody can separate a permanent reprice from a three-day promotion, and the discount history is gone for good.
  • Interpreting availability at capture time. "Ships in 2–3 weeks" becomes true in an in_stock boolean, and six months later nobody can reconstruct what the page actually said.
  • Trusting JSON-LD without verifying it. It is the most stable source on the page and also the one most often left stale after a promotion — check the sale price against the rendered page on a handful of listings before you rely on the field.
  • Keeping the price as a string with its symbol attached. €1.299,00 and $1,299.00 both parse to 1299, and both parse to 1.299 if you guess the separator wrong. Keep the number and the currency in separate columns.
  • Reading a market you did not mean to read. Several large retailers decide currency and price from a cookie rather than from the URL, so a run can capture the wrong country's figure correctly, with no error raised anywhere.

Check these before you scale up

Run one page per competitor and look at the output with your own eyes. Five minutes here saves a month.

  • A product that is on sale — do you get both prices, in the right columns?
  • A product that is out of stock — do you get a price at all, and is it the right one?
  • A multi-variant product — which variant did you capture, and did you mean to?
  • A product in a second currency or market domain, if the competitor has one
  • The most expensive product you track — decimal separator errors turn 1.299,00 into 1.29 and nothing downstream will question it

You now have competitor rows. The next lesson is the one that decides whether any of them mean anything.

Lesson 5: matching their listings to your catalogue

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.