Competitor price monitoring is usually described as "automatically checking what your rivals charge". That description is accurate and almost useless, because it hides where all the work is. Checking a price is the part a script does in two hundred milliseconds. Everything expensive happens either side of it.
A working pipeline has four stages. If you can name which stage is currently failing, you can fix it. If you cannot, you will keep rewriting the scraper and wondering why the numbers still look wrong.
The four stages
Read the right-hand column first. It is where the projects die.
| Stage | What it does | How it fails |
|---|---|---|
| 1. Selection | Decides which competitors and which of your products are worth watching | Someone picks "all of them", the cost triples and nobody reads the output |
| 2. Collection | Finds the competitor's product pages and reads price, stock and shipping off them | A redesign changes the page and the run returns fewer rows — quietly |
| 3. Matching | Decides that their listing and your SKU are the same sellable thing | A 2 litre bottle is compared with a 1 litre bottle and the margin report is fiction |
| 4. Action | Puts the number in front of the person or system that changes a price | The feed lands in a dashboard, the dashboard is opened twice, the project is cancelled |
Only stage two is scraping
This is the thing worth internalising before you spend any money. Stage two — collection — is a solved, commoditised problem. Dozens of vendors, including us, will do it. It is the stage every vendor demo spends its time on, because it is the stage that looks impressive on a screen.
Stages one, three and four are where your pipeline becomes useful or becomes shelfware, and none of them are technical. Stage one is a commercial judgement about where your margin is actually under threat. Stage three is a data problem that is never fully solved, only managed to an acceptable error rate. Stage four is an organisational problem about who is allowed to change a price and what evidence they need.
A vendor who only sells you stage two has sold you a quarter of a pipeline. That can be exactly right — plenty of teams genuinely only need the collection outsourced — but you should know which quarter you bought.
What it is not
It is not real-time. Almost every retail pricing decision is made on a daily or weekly cycle, by a human, in a meeting or a spreadsheet. Systems that promise minute-level price feeds are solving a marketplace buy-box problem, which is a genuinely different architecture and usually a genuinely different budget. If your prices change weekly, a daily feed is already faster than your process can absorb.
It is not a competitor's internal data. You see the shelf price a shopper sees. You do not see their cost, their margin, their stock depth, or the discount they are giving a key account. Treating the shelf price as a proxy for their strategy is the single most common analytical error in this field.
It is not a substitute for knowing your own catalogue. Most matching failures are not caused by the competitor's data being messy. They are caused by your own product data — missing EANs, pack sizes buried in free text, three internal SKUs that are the same physical item. Lesson five is blunt about this.
Signs you need one
Any two of these and the business case usually writes itself.
- Someone on your team has a browser folder of competitor product pages they open manually
- You have been undercut on a top-selling line and found out from a sales drop, not from a report
- Your repricing rules reference competitor prices that are updated "when we get round to it"
- You cannot currently answer "how many of our top 200 SKUs are above market right now?" in under an hour
- A supplier is enforcing minimum advertised prices and you have no evidence of who is breaking them
A note on cost shape
Collection is priced per page, almost everywhere, by almost everyone. That has one consequence worth planning around from the start: your bill is set by the number of competitor product pages times the number of times you look at them, and nothing else. Not by how many fields you extract. Not by how clever the matching is.
So the lever that controls your spend is lesson two, not lesson four. Twelve hundred pages checked daily is a very different monthly number from twelve thousand, and the second list is almost never twice as useful as the first. Teams that start with "track everything" and work down spend months paying for rows nobody reads.
Worked example: the page count behind a 1,200-SKU catalogue
Everything downstream — the bill, the run time, the argument with finance — falls out of one multiplication, and it is worth doing on paper before you speak to anyone selling you a tool. Take a retailer with 1,200 SKUs worth watching and five competitors on the shortlist. Read the last three rows together: the same five competitors and the same 1,200 products cost 108,000 pages a month or 38,600, depending on one scheduling decision.
| Step | Figure | Pages |
|---|---|---|
| SKUs you decide to watch | 1,200 | — |
| Competitors on the list | 5 | — |
| Share of your range each one actually carries | about 60% | 720 per competitor |
| One full pass over all five | 5 × 720 | 3,600 |
| Daily, across a month | 3,600 × 30 | 108,000 |
| Weekly, across a month | 3,600 × 4.3 | about 15,500 |
| Daily on the top 300, weekly on the remaining 900 | (5 × 180 × 30) + (5 × 540 × 4.3) | about 38,600 |
What usually goes wrong
Five failures account for most of the price feeds that get built and then quietly stop being opened.
- Budgeting for the catalogue instead of the overlap. Teams multiply 1,200 SKUs by five competitors, get 6,000 pages a run, and price the project off a figure two-thirds higher than reality. You only ever fetch pages that exist.
- Buying daily when the decision is weekly. If prices are reviewed in a Monday meeting, a daily run produces six snapshots nobody opens and multiplies the bill by seven.
- Starting at stage two. Collection is the stage with a demo attached, so it gets built first, and the matching problem in stage three is discovered after the contract is signed.
- Leaving stage four unowned. A feed with no named person who changes a price because of it is a reporting line item, not a pricing capability.
- Treating the pilot's page count as the steady-state count. Pilots run on 50 SKUs and one competitor. The multiplication above is what arrives in month two.
Next: the decision that sets both your cost and your usefulness — which competitors and which of your own products go on the list.
Lesson 2: choosing what to track