Now the scraping. This is the part everyone pictures when they imagine price monitoring, and it is genuinely the least interesting stage — which is good news, because it means you can treat it as plumbing and spend your attention elsewhere.
Two decisions matter here. What to extract, and how to make it survive the competitor redesigning their site.
Price is never one number
The most common beginner mistake is a single column called price. A product page typically shows a crossed-out original, a current selling price, sometimes a member price, sometimes a per-unit price underneath, and sometimes a basket-only price that is not displayed at all until checkout. Collapsing that into one number throws away the thing you most want to know, which is whether you are being undercut by a permanent reposition or by a three-day promotion.
Extract them separately and derive the comparison later. A discount percentage calculated at extraction time is a value you cannot recheck; two prices stored side by side can be re-derived any way you like, forever.
The field list
The first five are non-negotiable. The rest earn their place case by case.
| Field | Take it? | Note |
|---|---|---|
| price | Always | The current selling price as displayed, digits only, no symbol |
| list_price | Always | The crossed-out original where one is shown; empty is a valid answer |
| currency | Always | Stop guessing from the domain. Multi-market stores serve several currencies on one host. |
| availability | Always | The exact words on the page, not your interpretation of them |
| url | Always | The canonical URL, so you can click through and verify a surprising row |
| pack_size / unit | Usually | The single most valuable matching field in grocery, chemicals and anything sold by volume |
| shipping_cost | Sometimes | Often only visible at checkout. If it is on the page, take it; do not build a basket flow to get it. |
| seller_name | Marketplaces only | Meaningless on a single-brand store; essential on one where third parties list |
| ean / mpn | If shown | Worth more than every other field combined when it comes to matching |
| review_count | Rarely | Interesting, never actionable, and it changes every day so it bloats your change history |
Four places a price hides
Knowing which of these a site uses tells you immediately how fragile your extraction will be.
In the HTML, in a predictable element. The classic case. A CSS selector reads it, and it breaks the day the competitor changes their theme.
In JSON-LD structured data. Many stores publish a machine-readable Product block in the page source, for Google. When it is there and it is accurate, it is the best source available: it is explicitly a contract with search engines, so it changes far less often than the visual markup. It is also frequently incomplete, wrong on sale prices, or absent on exactly the pages you care about — so verify, do not assume.
In an internal API call the page makes after loading. The HTML arrives with a price-shaped hole in it and JavaScript fills it in. Any extraction that reads raw HTML gets nothing; you need something that runs the page's JavaScript first.
In an image. Rare, deliberate, and a signal that the site does not want to be read. Treat it as a competitor you track manually.
Selectors versus a model that reads the page
A CSS selector is precise, free, fast and brittle. It encodes the competitor's current HTML structure into your pipeline, and it breaks on redesign — not with an error, which would be helpful, but usually by returning nothing or by returning the wrong element that happens to sit where the old one did.
The alternative is to describe the field in words and have a model find it on the rendered page. That survives a redesign, because a price is still visibly a price after the CSS changes. It costs more per page and it is not infallible — a model can also pick the wrong element, and it will do so with complete confidence.
In practice the deciding factor is maintenance headcount. If nobody owns fixing broken selectors within a day, selectors are a false economy: the cheapest extraction in the world is worthless on the mornings it returns nothing.
Worked example: one product page, eight columns
A single listing, as a shopper sees it on the left and as it should land in your table on the right. Every judgement has been deferred: nothing here is interpreted, only recorded. Note what is absent — there is no discount column, because 1 − 199 ÷ 249 is 20.1% today and will still be 20.1% whenever anyone asks, recomputed from two columns that were written down.
| What the page shows | Column | Stored value |
|---|---|---|
| €249.00, struck through | list_price | 249.00 |
| €199.00, in red | sale_price | 199.00 |
| The € symbol before the figure | currency | EUR |
| "In stock — 3 left" | availability_text | In stock — 3 left |
| nothing on the page | availability_class | left empty, derived downstream |
| "Free delivery over €50" | shipping_text | Free delivery over €50 |
| read at 06:14 on 14 March | captured_at | 2026-03-14T06:14:00Z |
| the page the figures came from | source_url | https://…/p/12345 |
What usually goes wrong
Extraction failures are rarely loud. Four of these five produce a table that looks entirely reasonable and is wrong.
- One price column. The page showed two numbers and the table kept one, so nobody can separate a permanent reprice from a three-day promotion, and the discount history is gone for good.
- Interpreting availability at capture time. "Ships in 2–3 weeks" becomes true in an in_stock boolean, and six months later nobody can reconstruct what the page actually said.
- Trusting JSON-LD without verifying it. It is the most stable source on the page and also the one most often left stale after a promotion — check the sale price against the rendered page on a handful of listings before you rely on the field.
- Keeping the price as a string with its symbol attached. €1.299,00 and $1,299.00 both parse to 1299, and both parse to 1.299 if you guess the separator wrong. Keep the number and the currency in separate columns.
- Reading a market you did not mean to read. Several large retailers decide currency and price from a cookie rather than from the URL, so a run can capture the wrong country's figure correctly, with no error raised anywhere.
Check these before you scale up
Run one page per competitor and look at the output with your own eyes. Five minutes here saves a month.
- A product that is on sale — do you get both prices, in the right columns?
- A product that is out of stock — do you get a price at all, and is it the right one?
- A multi-variant product — which variant did you capture, and did you mean to?
- A product in a second currency or market domain, if the competitor has one
- The most expensive product you track — decimal separator errors turn 1.299,00 into 1.29 and nothing downstream will question it
You now have competitor rows. The next lesson is the one that decides whether any of them mean anything.
Lesson 5: matching their listings to your catalogue