[{"data":1,"prerenderedAt":1288},["ShallowReactive",2],{"learn-course-competitor-price-monitoring":3,"learn-courses":1086},{"slug":4,"order":5,"level":6,"time":7,"card_text":8,"seo":9,"hero":15,"outcomes":25,"who":35,"syllabus":46,"faq":49,"lessons":66},"competitor-price-monitoring",1,"No coding required","8 lessons, about 90 minutes","From \"we check three competitors by hand on Mondays\" to a feed you trust enough to reprice from. The eight decisions in order, including the two that quietly ruin most projects.",{"title":10,"description":11,"keywords":12,"og_title":13,"og_description":14},"Competitor Price Monitoring: A Free 8-Lesson Course","Build a competitor price monitoring pipeline end to end: pick competitors, find product URLs, extract prices, match to your catalogue, schedule, export, reprice. Free, ungated.","competitor price monitoring, price monitoring course, how to track competitor prices, price tracking pipeline, competitor price tracking tutorial, retail price monitoring","A free course on building a competitor price monitoring pipeline","Eight written lessons on the part nobody covers: matching, data quality, and turning prices into decisions.",{"badge":16,"title":17,"subtitle":18,"cta_primary":19,"cta_secondary":22},"Course one","Build a competitor price monitoring pipeline","Price monitoring looks like a scraping problem for about a week. Then you discover that scraping was the easy part, and the project actually lives or dies on which competitors you picked, whether their listings are really the same product as yours, and whether anyone notices the morning the feed comes back half empty. This course is those eight decisions, in the order you have to make them.",{"label":20,"url":21},"Start with lesson one","/learn/competitor-price-monitoring/what-is-competitor-price-monitoring",{"label":23,"url":24},"Try the free price checker","/tools/free-competitor-price-checker",{"title":26,"items":27},"What you will be able to do",[28,29,30,31,32,33,34],"Say which competitors matter to your margin, and defend the list to a finance director","Collect competitor product URLs without copying them one at a time","Decide which fields you actually need, and which ones are a trap","Judge a match rate honestly, including the denominator that most vendors quietly change","Spot a feed that has silently degraded before a buyer acts on it","Put the numbers where the decision is made, instead of in a dashboard nobody opens","Write a repricing rule that survives contact with a competitor who is out of stock",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Pricing, category and e-commerce managers who own the number","Marketplace sellers repricing against a handful of named rivals","Anyone evaluating price monitoring vendors who wants to ask better questions","Not written for",[44,45],"Engineers who want a tutorial on CSS selectors and headless browsers — that is a different and much shorter problem","Anyone looking for a sub-second price feed; nothing here runs faster than daily",{"title":47,"intro":48},"The eight lessons","Each one ends where the next one starts. Four lessons are pure method; the ones that walk through our product say so in their own heading.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","Questions people ask in the first ten minutes.",[54,57,60,63],{"title":55,"description":56},"Do I need a Scrapewise account to follow this?","For five of the eight lessons, no. Lessons three, four and seven walk through doing the thing in our portal and are marked as needing an account. A new account comes with five free requests and no card, which covers those lessons.",{"title":58,"description":59},"Is scraping competitor prices legal?","This is not legal advice. Reading a publicly visible price is not the same question as storing it, republishing it, or building a product on it, and the answer changes by jurisdiction and by the site's own terms. Lesson two covers how to keep the collection narrow enough that the question stays simple, but take your own advice before you start.",{"title":61,"description":62},"How often can prices be checked?","This course assumes a daily cycle, which is what most retail pricing decisions actually run on. If you need minute-level reaction — a marketplace buy box, an airline fare — the architecture in lesson six is wrong for you and you want an event-driven system instead.",{"title":64,"description":65},"What if my competitors are not online retailers?","The method holds for any public price list: distributors, wholesalers, rental catalogues. What changes is lesson three, because those sites rarely have a clean product sitemap. Lesson three covers that case directly.",[67,194,332,449,599,741,859,974],{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":185,"next_step":190},"what-is-competitor-price-monitoring","What it actually is","What competitor price monitoring actually is","The four stages of a price pipeline, why only two of them are scraping, and the one question to ask before you build anything.","9 min",false,{"title":75,"description":76,"keywords":77},"What Is Competitor Price Monitoring? The Four Stages","Competitor price monitoring is four stages, and only two involve scraping. Here is what each stage does, where projects fail, and how to tell if you need one at all.","what is competitor price monitoring, price monitoring definition, competitor price tracking explained, price intelligence stages",[79,84,109,115,120,126,136,141,176],{"type":80,"paragraphs":81},"prose",[82,83],"Competitor price monitoring is usually described as \"automatically checking what your rivals charge\". That description is accurate and almost useless, because it hides where all the work is. Checking a price is the part a script does in two hundred milliseconds. Everything expensive happens either side of it.","A working pipeline has four stages. If you can name which stage is currently failing, you can fix it. If you cannot, you will keep rewriting the scraper and wondering why the numbers still look wrong.",{"type":85,"title":86,"intro":87,"headers":88,"rows":92},"table","The four stages","Read the right-hand column first. It is where the projects die.",[89,90,91],"Stage","What it does","How it fails",[93,97,101,105],[94,95,96],"1. Selection","Decides which competitors and which of your products are worth watching","Someone picks \"all of them\", the cost triples and nobody reads the output",[98,99,100],"2. Collection","Finds the competitor's product pages and reads price, stock and shipping off them","A redesign changes the page and the run returns fewer rows — quietly",[102,103,104],"3. Matching","Decides that their listing and your SKU are the same sellable thing","A 2 litre bottle is compared with a 1 litre bottle and the margin report is fiction",[106,107,108],"4. Action","Puts the number in front of the person or system that changes a price","The feed lands in a dashboard, the dashboard is opened twice, the project is cancelled",{"type":80,"title":110,"paragraphs":111},"Only stage two is scraping",[112,113,114],"This is the thing worth internalising before you spend any money. Stage two — collection — is a solved, commoditised problem. Dozens of vendors, including us, will do it. It is the stage every vendor demo spends its time on, because it is the stage that looks impressive on a screen.","Stages one, three and four are where your pipeline becomes useful or becomes shelfware, and none of them are technical. Stage one is a commercial judgement about where your margin is actually under threat. Stage three is a data problem that is never fully solved, only managed to an acceptable error rate. Stage four is an organisational problem about who is allowed to change a price and what evidence they need.","A vendor who only sells you stage two has sold you a quarter of a pipeline. That can be exactly right — plenty of teams genuinely only need the collection outsourced — but you should know which quarter you bought.",{"type":116,"variant":117,"title":118,"text":119},"callout","note","The question to ask before building anything","\"If I had perfect competitor prices tomorrow morning, what would change by Friday?\" If the honest answer is \"we would look at them\", you do not have a monitoring problem yet, you have a pricing-process problem. Fix the decision first; the data is cheap once something is waiting for it.",{"type":80,"title":121,"paragraphs":122},"What it is not",[123,124,125],"It is not real-time. Almost every retail pricing decision is made on a daily or weekly cycle, by a human, in a meeting or a spreadsheet. Systems that promise minute-level price feeds are solving a marketplace buy-box problem, which is a genuinely different architecture and usually a genuinely different budget. If your prices change weekly, a daily feed is already faster than your process can absorb.","It is not a competitor's internal data. You see the shelf price a shopper sees. You do not see their cost, their margin, their stock depth, or the discount they are giving a key account. Treating the shelf price as a proxy for their strategy is the single most common analytical error in this field.","It is not a substitute for knowing your own catalogue. Most matching failures are not caused by the competitor's data being messy. They are caused by your own product data — missing EANs, pack sizes buried in free text, three internal SKUs that are the same physical item. Lesson five is blunt about this.",{"type":127,"title":128,"intro":129,"items":130},"list","Signs you need one","Any two of these and the business case usually writes itself.",[131,132,133,134,135],"Someone on your team has a browser folder of competitor product pages they open manually","You have been undercut on a top-selling line and found out from a sales drop, not from a report","Your repricing rules reference competitor prices that are updated \"when we get round to it\"","You cannot currently answer \"how many of our top 200 SKUs are above market right now?\" in under an hour","A supplier is enforcing minimum advertised prices and you have no evidence of who is breaking them",{"type":80,"title":137,"paragraphs":138},"A note on cost shape",[139,140],"Collection is priced per page, almost everywhere, by almost everyone. That has one consequence worth planning around from the start: your bill is set by the number of competitor product pages times the number of times you look at them, and nothing else. Not by how many fields you extract. Not by how clever the matching is.","So the lever that controls your spend is lesson two, not lesson four. Twelve hundred pages checked daily is a very different monthly number from twelve thousand, and the second list is almost never twice as useful as the first. Teams that start with \"track everything\" and work down spend months paying for rows nobody reads.",{"type":85,"title":142,"intro":143,"headers":144,"rows":148},"Worked example: the page count behind a 1,200-SKU catalogue","Everything downstream — the bill, the run time, the argument with finance — falls out of one multiplication, and it is worth doing on paper before you speak to anyone selling you a tool. Take a retailer with 1,200 SKUs worth watching and five competitors on the shortlist. Read the last three rows together: the same five competitors and the same 1,200 products cost 108,000 pages a month or 38,600, depending on one scheduling decision.",[145,146,147],"Step","Figure","Pages",[149,153,156,160,164,168,172],[150,151,152],"SKUs you decide to watch","1,200","—",[154,155,152],"Competitors on the list","5",[157,158,159],"Share of your range each one actually carries","about 60%","720 per competitor",[161,162,163],"One full pass over all five","5 × 720","3,600",[165,166,167],"Daily, across a month","3,600 × 30","108,000",[169,170,171],"Weekly, across a month","3,600 × 4.3","about 15,500",[173,174,175],"Daily on the top 300, weekly on the remaining 900","(5 × 180 × 30) + (5 × 540 × 4.3)","about 38,600",{"type":127,"title":177,"intro":178,"items":179},"What usually goes wrong","Five failures account for most of the price feeds that get built and then quietly stop being opened.",[180,181,182,183,184],"Budgeting for the catalogue instead of the overlap. Teams multiply 1,200 SKUs by five competitors, get 6,000 pages a run, and price the project off a figure two-thirds higher than reality. You only ever fetch pages that exist.","Buying daily when the decision is weekly. If prices are reviewed in a Monday meeting, a daily run produces six snapshots nobody opens and multiplies the bill by seven.","Starting at stage two. Collection is the stage with a demo attached, so it gets built first, and the matching problem in stage three is discovered after the contract is signed.","Leaving stage four unowned. A feed with no named person who changes a price because of it is a reporting line item, not a pricing capability.","Treating the pilot's page count as the steady-state count. Pilots run on 50 SKUs and one competitor. The multiplication above is what arrives in month two.",[186,187,188,189],"A price pipeline is four stages: selection, collection, matching, action. Only collection is scraping.","Vendors demo stage two because it looks good. Your project will fail or succeed on one, three and four.","If nothing downstream is waiting for the data, you have a process problem, not a data problem.","Your bill is pages times frequency, so the scope decision in lesson two is also the budget decision.",{"text":191,"label":192,"url":193},"Next: the decision that sets both your cost and your usefulness — which competitors and which of your own products go on the list.","Lesson 2: choosing what to track","/learn/competitor-price-monitoring/choose-competitors-and-skus",{"slug":195,"nav_title":196,"title":197,"summary":198,"time":199,"needs_account":73,"seo":200,"blocks":204,"takeaways":323,"next_step":328},"choose-competitors-and-skus","Choosing what to track","Choosing which competitors and which SKUs to track","How to build a list that is small enough to afford and large enough to matter, using margin at risk rather than gut feel.","11 min",{"title":201,"description":202,"keywords":203},"Which Competitors and SKUs Should You Price Monitor?","A method for choosing what to track: rank competitors by overlap and threat, rank SKUs by margin at risk, and cut the list before it cuts your budget.","which competitors to track, price monitoring scope, competitor selection pricing, which skus to price monitor, margin at risk",[205,209,226,231,263,267,273,307,315],{"type":80,"paragraphs":206},[207,208],"Every price monitoring project starts with somebody saying \"let's track everyone\". It is a reasonable instinct and it is the main reason these projects get cancelled at renewal. A list that is ten times too long costs ten times too much, takes ten times as long to match, and produces a report with so many rows that the person who asked for it stops opening it.","The goal of this lesson is a defensible list: a set of competitors and a set of your own products where you can explain, line by line, why each one is there.",{"type":210,"title":211,"intro":212,"items":213},"steps","Ranking competitors","Score each candidate on two axes. You want the ones that are high on both.",[214,217,220,223],{"title":215,"text":216},"Overlap: how much of your range do they actually carry?","Take your top 100 products by revenue and check, by hand, how many of them the competitor stocks. Twenty minutes of clicking gives you a number. A competitor carrying 8 of your top 100 is not a competitor for pricing purposes, however loudly sales talks about them. A competitor carrying 70 of them sets your market price whether you like it or not.",{"title":218,"text":219},"Threat: does a shopper comparing you to them actually switch?","Price only matters where the alternative is credible. A rival with worse delivery, no stock, or a market the customer will not buy from can be cheaper than you all year and cost you nothing. The honest test is whether your own sales team mentions them in lost-deal notes. If they never come up, they are a benchmark, not a threat — and benchmarks can be checked monthly instead of daily.",{"title":221,"text":222},"Separate marketplaces from retailers","A marketplace listing is not one competitor, it is a rotating cast of sellers behind one URL, and the price you capture is whoever won the buy box at the moment you looked. That is still useful, but treat it as a market signal rather than as a named rival, and do not build a repricing rule that chases it directly.",{"title":224,"text":225},"Cut to between three and eight","Almost every retailer we have worked with ends up with a daily list in that range, no matter how big they are, plus a longer tail checked weekly or monthly. If your daily list has twenty names on it, two things are true: you cannot name what you would do differently based on fifteen of them, and you are paying for all twenty.",{"type":80,"title":227,"paragraphs":228},"Ranking your own products",[229,230],"The competitor list is the easy half. The expensive half is deciding which of your own SKUs are in scope, and here the useful metric is margin at risk, not revenue.","Margin at risk is roughly: units sold, times your margin per unit, times how price-sensitive that line is. A high-revenue product where you are the only stockist in the country has almost no margin at risk — nobody is going to undercut you. A mid-revenue product sold by four rivals within two percent of each other has enormous margin at risk, because a single competitor move changes your conversion rate that week.",{"type":85,"title":232,"intro":233,"headers":234,"rows":239},"A worked cut","A mid-size electronics retailer, roughly 9,000 SKUs. This is the shape of the exercise, not a template to copy.",[235,236,237,238],"Segment","SKUs","In scope?","Why",[240,245,250,255,259],[241,242,243,244],"Top sellers, 3+ rivals stock them","180","Daily","This is the margin at risk. Everything else is a rounding error against it.",[246,247,248,249],"Top sellers, sole stockist","60","Monthly","Nobody to be undercut by. Checked only to catch a new entrant.",[251,252,253,254],"Long tail, any rival","6,200","No","Individually immaterial; collectively would be 90% of the bill.",[256,257,243,258],"New lines, first 90 days","~120 rotating","Launch pricing is where mistakes are cheapest to catch and most expensive to leave.",[260,261,253,262],"Own-brand / exclusive","2,400","No comparable listing exists, so there is nothing to match against.",{"type":116,"variant":264,"title":265,"text":266},"warning","The own-brand trap","Exclusive and own-brand lines are the single biggest source of wasted scope. They look like normal products in your catalogue, they get included by default when someone exports \"all active SKUs\", and they can never be matched to anything because no competitor sells them. They inflate your page count and then depress your match rate, so the project looks like it is failing at stage three when it is actually failing at stage one.",{"type":80,"title":268,"paragraphs":269},"Do the arithmetic before you commit",[270,271,272],"Your daily page count is competitors times in-scope SKUs that each competitor actually carries. Not competitors times your whole catalogue — a rival who stocks 40% of your in-scope range only contributes 40% of those pages.","For the example above: 180 daily SKUs, five competitors, average carriage of about 60%, gives roughly 540 pages a day, or about 16,000 a month. That is a number you can take to whoever signs off the budget, and it is small enough that you can afford to check it twice a day later if the business case justifies it.","Now run the same sum for the \"track everything\" version: 9,000 SKUs across five competitors is in the region of 800,000 pages a month. Same project, fifty times the bill, and a report nobody can read.",{"type":85,"title":274,"intro":275,"headers":276,"rows":282},"Worked example: ranking four lines by margin at risk","Margin at risk is monthly margin times the share of it a cheaper competitor could plausibly take. Revenue puts these four lines in the order A, C, D, B. Margin at risk puts them in the order C, A, D, B — promoting the second-biggest line by revenue to the top, and demoting the fattest-margin product in the business to nothing at all, because nobody else sells it.",[277,278,279,280,281],"Line","Monthly revenue","Monthly margin","Share a cheaper rival could take","Margin at risk",[283,289,295,301],[284,285,286,287,288],"A. Branded 55-inch TV, every rival stocks it","€180,000","€7,200","90%","€6,480",[290,291,292,293,294],"B. Own-brand HDMI cable","€36,000","€24,000","0%","€0",[296,297,298,299,300],"C. Branded coffee machine, three rivals stock it","€110,000","€20,000","70%","€14,000",[302,303,304,305,306],"D. Sole-stockist power tool","€60,000","€18,000","10%","€1,800",{"type":127,"title":177,"intro":308,"items":309},"The selection stage has no error messages, so these are only ever discovered through the bill or through a report nobody trusts.",[310,311,312,313,314],"Scoring competitors on how often the sales team mentions them. Irritation is not overlap. Twenty minutes of clicking through your top 100 is the only input here that survives scrutiny.","Letting the list grow by one competitor per meeting. Every name added is a permanent multiplier on every future run, and nobody ever proposes removing one.","Tracking own-brand lines because they are top sellers. They cannot be matched to anything, so they generate pages, cost money and return nothing — line B above is the whole argument.","Treating one marketplace as one competitor. A single listing page can carry fifteen sellers, and the price that matters is whoever currently holds the buy box, not the marketplace's own.","Pricing the project on the SKU count rather than the SKU-times-competitor count, and discovering the real figure after the budget is signed off.",{"type":127,"title":316,"intro":317,"items":318},"Write these down before you move on","Lesson three assumes you have them.",[319,320,321,322],"The named competitor list, split into daily and less-than-daily","The in-scope SKU list, with the rule you used to build it written next to it","Your estimated monthly page count, and therefore your estimated monthly cost","The review date — scope rots, and a list nobody has revisited in a year is tracking last year's market",[324,325,326,327],"Score competitors on overlap with your range and on whether a shopper genuinely switches. Most daily lists should be three to eight names.","Rank your own products by margin at risk, not revenue. Sole-stockist lines have almost none.","Own-brand and exclusive SKUs can never be matched — excluding them is the cheapest single improvement to your match rate.","Page count equals competitors times the SKUs they actually carry. Do that arithmetic before you sign anything.",{"text":329,"label":330,"url":331},"You have a list of competitors and a list of products. Next you need the thing that connects them: the actual URLs of their product pages.","Lesson 3: finding competitor product URLs","/learn/competitor-price-monitoring/find-competitor-product-urls",{"slug":333,"nav_title":334,"title":335,"summary":336,"time":337,"needs_account":338,"seo":339,"blocks":343,"takeaways":440,"next_step":445},"find-competitor-product-urls","Finding product URLs","Finding every competitor product URL without copying them by hand","Four ways to get a competitor's full product URL list, ranked by how much work they are, and what to do when none of them work.","12 min",true,{"title":340,"description":341,"keywords":342},"How to Get a Competitor's Full Product URL List","Four methods for collecting every product URL on a competitor site — sitemaps, category crawling, search, and a URL extractor — plus what to do when the site hides them.","find competitor product urls, extract product urls, product url list, scrape product links, xml sitemap products, competitor catalogue",[344,348,363,370,377,405,408,427,435],{"type":80,"paragraphs":345},[346,347],"A scraper needs addresses. Before anything can read a price, something has to produce a list of competitor product pages — ideally the complete list, ideally kept up to date as they add and drop lines.","This is the step teams most often do by hand, and it is the step where doing it by hand hurts most: a thousand URLs copied from a browser is a week of somebody's life, and it is stale by the time it is finished. There are four ways to avoid that, and they are worth trying in this order.",{"type":210,"title":349,"items":350},"The four methods, easiest first",[351,354,357,360],{"title":352,"text":353},"1. The XML sitemap","Most e-commerce platforms publish one, because Google needs it. Try /sitemap.xml and /robots.txt on the competitor's domain; robots.txt usually names the sitemap even when the default path is wrong. What you get is often a sitemap index pointing at a dozen child files, one of which is products. This is the best possible outcome: a complete, maintained, machine-readable list of every product page, published deliberately. On a well-run store it takes about five minutes.",{"title":355,"text":356},"2. Category pages","When there is no product sitemap, walk the category tree instead: open a listing page, take every product link on it, follow the pagination, repeat. Slower and noisier than a sitemap — you will pick up promotional tiles and cross-sell blocks — but it works on nearly every store, and it has one genuine advantage: it tells you which category the competitor files a product under, which is useful context for matching later.",{"title":358,"text":359},"3. Their own site search","If you only care about a few hundred known products, searching the competitor's site for each EAN or manufacturer part number is often the fastest route, and it gives you something the other methods do not: a direct, high-confidence link between their URL and your SKU. That link is worth a great deal in lesson five. The limitation is obvious — it only finds what you already know to look for.",{"title":361,"text":362},"4. A URL extractor","A tool that takes the domain and returns the product URLs, handling the sitemap-or-crawl decision for you. We publish a free one; so do others. This is method one and two with the plumbing hidden, which is worth it when you are doing this for eight competitors rather than one.",{"type":116,"variant":364,"title":365,"text":366,"cta":367},"product","Doing it with our free tool","The Competitor Product URL Extractor takes a store URL and returns the product page addresses it can find, with no account and no card. It is the fastest way to find out whether a given competitor is going to be easy or hard before you commit to tracking them. Note the honest limitation: stores that publish no sitemap and render their listings entirely in JavaScript can come back with nothing.",{"label":368,"url":369},"Open the free URL extractor","/tools/free-product-url-extractor",{"type":80,"title":371,"paragraphs":372},"Cleaning the list",[373,374,375,376],"Whatever method produced it, the raw list is never the list you want. Three passes fix most of it.","First, strip tracking parameters. A URL ending in ?utm_source=newsletter is the same page as the one without it, but a naive pipeline will treat them as two pages and bill you for both. Cut everything after the question mark unless you know a parameter is load-bearing — on some stores the variant selector lives there, and you will see it because the page genuinely changes when you remove it.","Second, collapse variants. A t-shirt in six sizes is often six URLs with the same price. Decide deliberately whether you are tracking the product or the variant, because tracking all six multiplies your bill by six and usually tells you one thing. Where sizes really are priced differently — shoes, tyres, anything sold by capacity — you do want them separately, and you should say so explicitly rather than letting the extractor decide.","Third, drop what is not a product. Category pages, blog posts, gift cards, bundles, store locators. They sneak in from every method and each one is a page you pay to read and a row that will never match.",{"type":85,"title":378,"intro":379,"headers":380,"rows":384},"Choosing a method per competitor","In practice you will use different methods for different sites, and that is fine.",[381,382,383],"Situation","Use","Watch out for",[385,389,393,397,401],[386,387,388],"Product sitemap exists and is current","Sitemap","Sitemaps that include discontinued lines for months",[390,391,392],"No sitemap, normal server-rendered listings","Category crawl","Infinite scroll with no page numbers",[394,395,396],"You only track 200 known EANs","Their site search","Zero-result pages returning a 200 status",[398,399,400],"Eight competitors, limited patience","URL extractor","JavaScript-only storefronts returning nothing",[402,403,404],"Login required to see prices","None of these","Do not. A price behind a login is not a public price.",{"type":116,"variant":264,"title":406,"text":407},"Stop at the login wall","If a competitor's prices are only visible to registered trade customers, the whole exercise changes character — from reading a public page to accessing an account under terms somebody agreed to. This course does not cover that, and you should talk to whoever handles your legal questions before anyone on your team creates that account.",{"type":210,"title":409,"intro":410,"items":411},"Worked example: one competitor's sitemap to a payable URL list","Start to finish on a single competitor, with the figure each step removes. The absolute counts will differ for your sites; the shape of the reduction will not.",[412,415,418,421,424],{"title":413,"text":414},"Fetch /sitemap.xml and follow the index","Most sitemaps are an index pointing at several child files. You want only the product ones — if they are named products-1.xml through products-4.xml, take all four and none of the blog, category or store-locator files. Say the four together list 41,000 URLs.",{"title":416,"text":417},"Strip query strings and fragments, then deduplicate","Remove everything after ? and #. Tracking parameters, sort orders and session ids turn one page into several, and every duplicate is a page you pay to fetch twice. Say this leaves 38,600 distinct URLs.",{"title":419,"text":420},"Keep only the paths that match the product pattern","Open three product pages by hand and find the shared segment — /p/, /product/, /dp/, or a trailing numeric id. Everything that does not match is a category, a filter permutation or an article. Say 29,000 survive.",{"title":422,"text":423},"Decide products or variants, and write the decision down","If the sitemap lists every colour and size separately, those 29,000 URLs may be 9,000 products. Collapsing to the parent cuts the fetch count by roughly two-thirds and loses per-variant stock. Collapsing is usually right for pricing and usually wrong for availability, which is why the decision belongs in writing rather than in whoever built it.",{"title":425,"text":426},"Intersect with the SKUs you already chose","You picked 1,200 SKUs in lesson two. Only the overlap is worth fetching. If 700 of your 1,200 appear in their list, this competitor costs 700 pages a run, not 29,000 — and that single intersection is the difference between a project and a quote you walk away from.",{"type":127,"title":177,"intro":428,"items":429},"Discovery is the cheapest stage to get right and the most expensive to get wrong, because the error multiplies by every future run.",[430,431,432,433,434],"Taking the whole sitemap. It is the single most expensive mistake in this lesson: 29,000 pages a run instead of 700, every run, forever.","Missing the sitemap index and reading only the first child file, then wondering why a third of the catalogue never appears.","Leaving tracking parameters attached. The same product arrives three times under three URLs, and the deduplication problem moves downstream into matching, where it is far harder to see.","Discovering URLs once and never again. Competitors add and drop lines constantly, so a list built in January is measurably wrong by March and silently wrong the entire time.","Pushing past a login wall because the data behind it is better. That is the one method on this list with a legal answer rather than a technical one, and course six is where it gets answered.",{"type":80,"title":436,"paragraphs":437},"Keeping the list fresh",[438,439],"Catalogues move. A competitor adds forty lines before Christmas and drops a hundred in January, and a URL list captured once decays at maybe one to three percent a month depending on the category. The symptom is subtle: your page count stays flat while the number of pages that return a real price slowly falls, and your coverage quietly erodes.","The fix is boring and effective: re-run whichever method you used on a schedule — monthly is usually enough — and diff the result against the current list. New URLs get added, URLs that have disappeared from the sitemap for two consecutive runs get retired. That diff is also a small, genuinely interesting competitive signal in its own right: it is a list of what your rival started and stopped selling.",[441,442,443,444],"Try the XML sitemap first. It is complete, maintained and takes five minutes.","Strip tracking parameters and decide explicitly whether you track products or variants — both affect the bill directly.","Different competitors warrant different methods; there is no need to pick one for all of them.","Re-run the discovery monthly and diff it. The diff doubles as a record of what your competitor launched and dropped.",{"text":446,"label":447,"url":448},"With addresses in hand, the next question is what to read off each page — and which tempting fields are a trap.","Lesson 4: extracting price, stock and shipping","/learn/competitor-price-monitoring/extract-price-stock-and-shipping",{"slug":450,"nav_title":451,"title":452,"summary":453,"time":337,"needs_account":338,"seo":454,"blocks":458,"takeaways":590,"next_step":595},"extract-price-stock-and-shipping","Extracting the fields","Getting price, stock and shipping off the page","Which fields to extract, why the sale price is two fields and not one, and the four ways a price appears on a page.",{"title":455,"description":456,"keywords":457},"Which Fields to Extract From a Competitor Product Page","Price is never one number. How to extract list price, sale price, currency, stock and shipping reliably, and which fields cost more than they are worth.","extract price from website, product page fields, scrape price and stock, json-ld price, competitor price fields, structured data product",[459,463,468,512,515,523,529,535,573,581],{"type":80,"paragraphs":460},[461,462],"Now the scraping. This is the part everyone pictures when they imagine price monitoring, and it is genuinely the least interesting stage — which is good news, because it means you can treat it as plumbing and spend your attention elsewhere.","Two decisions matter here. What to extract, and how to make it survive the competitor redesigning their site.",{"type":80,"title":464,"paragraphs":465},"Price is never one number",[466,467],"The most common beginner mistake is a single column called price. A product page typically shows a crossed-out original, a current selling price, sometimes a member price, sometimes a per-unit price underneath, and sometimes a basket-only price that is not displayed at all until checkout. Collapsing that into one number throws away the thing you most want to know, which is whether you are being undercut by a permanent reposition or by a three-day promotion.","Extract them separately and derive the comparison later. A discount percentage calculated at extraction time is a value you cannot recheck; two prices stored side by side can be re-derived any way you like, forever.",{"type":85,"title":469,"intro":470,"headers":471,"rows":475},"The field list","The first five are non-negotiable. The rest earn their place case by case.",[472,473,474],"Field","Take it?","Note",[476,480,483,486,489,492,496,500,504,508],[477,478,479],"price","Always","The current selling price as displayed, digits only, no symbol",[481,478,482],"list_price","The crossed-out original where one is shown; empty is a valid answer",[484,478,485],"currency","Stop guessing from the domain. Multi-market stores serve several currencies on one host.",[487,478,488],"availability","The exact words on the page, not your interpretation of them",[490,478,491],"url","The canonical URL, so you can click through and verify a surprising row",[493,494,495],"pack_size / unit","Usually","The single most valuable matching field in grocery, chemicals and anything sold by volume",[497,498,499],"shipping_cost","Sometimes","Often only visible at checkout. If it is on the page, take it; do not build a basket flow to get it.",[501,502,503],"seller_name","Marketplaces only","Meaningless on a single-brand store; essential on one where third parties list",[505,506,507],"ean / mpn","If shown","Worth more than every other field combined when it comes to matching",[509,510,511],"review_count","Rarely","Interesting, never actionable, and it changes every day so it bloats your change history",{"type":116,"variant":264,"title":513,"text":514},"Store the words, not your reading of them","\"Only 2 left\", \"Ships in 3-5 days\", \"Available to order\" and \"In stock\" are four different commercial situations. A pipeline that maps all of them to a boolean in_stock = true has destroyed information it cannot get back. Capture the exact string; classify it downstream where you can change your mind.",{"type":80,"title":516,"paragraphs":517},"Four places a price hides",[518,519,520,521,522],"Knowing which of these a site uses tells you immediately how fragile your extraction will be.","In the HTML, in a predictable element. The classic case. A CSS selector reads it, and it breaks the day the competitor changes their theme.","In JSON-LD structured data. Many stores publish a machine-readable Product block in the page source, for Google. When it is there and it is accurate, it is the best source available: it is explicitly a contract with search engines, so it changes far less often than the visual markup. It is also frequently incomplete, wrong on sale prices, or absent on exactly the pages you care about — so verify, do not assume.","In an internal API call the page makes after loading. The HTML arrives with a price-shaped hole in it and JavaScript fills it in. Any extraction that reads raw HTML gets nothing; you need something that runs the page's JavaScript first.","In an image. Rare, deliberate, and a signal that the site does not want to be read. Treat it as a competitor you track manually.",{"type":80,"title":524,"paragraphs":525},"Selectors versus a model that reads the page",[526,527,528],"A CSS selector is precise, free, fast and brittle. It encodes the competitor's current HTML structure into your pipeline, and it breaks on redesign — not with an error, which would be helpful, but usually by returning nothing or by returning the wrong element that happens to sit where the old one did.","The alternative is to describe the field in words and have a model find it on the rendered page. That survives a redesign, because a price is still visibly a price after the CSS changes. It costs more per page and it is not infallible — a model can also pick the wrong element, and it will do so with complete confidence.","In practice the deciding factor is maintenance headcount. If nobody owns fixing broken selectors within a day, selectors are a false economy: the cheapest extraction in the world is worthless on the mornings it returns nothing.",{"type":116,"variant":364,"title":530,"text":531,"cta":532},"Doing it in Scrapewise","You define the columns once in a Custom Schema — a field name, a type (text, number or yes/no), and a description that tells the extractor what to look for on the page. The description is the instruction, not a comment: \"stock status\" produces guesswork, while \"the exact text on the page showing stock status, for example In Stock, Out of Stock, Only 2 left\" produces the string you wanted. The same schema is then reusable across every competitor you point it at, which is what makes their columns line up.",{"label":533,"url":534},"See how custom scrapers work","/custom-scrapers",{"type":85,"title":536,"intro":537,"headers":538,"rows":542},"Worked example: one product page, eight columns","A single listing, as a shopper sees it on the left and as it should land in your table on the right. Every judgement has been deferred: nothing here is interpreted, only recorded. Note what is absent — there is no discount column, because 1 − 199 ÷ 249 is 20.1% today and will still be 20.1% whenever anyone asks, recomputed from two columns that were written down.",[539,540,541],"What the page shows","Column","Stored value",[543,546,550,553,557,561,565,569],[544,481,545],"€249.00, struck through","249.00",[547,548,549],"€199.00, in red","sale_price","199.00",[551,484,552],"The € symbol before the figure","EUR",[554,555,556],"\"In stock — 3 left\"","availability_text","In stock — 3 left",[558,559,560],"nothing on the page","availability_class","left empty, derived downstream",[562,563,564],"\"Free delivery over €50\"","shipping_text","Free delivery over €50",[566,567,568],"read at 06:14 on 14 March","captured_at","2026-03-14T06:14:00Z",[570,571,572],"the page the figures came from","source_url","https://…/p/12345",{"type":127,"title":177,"intro":574,"items":575},"Extraction failures are rarely loud. Four of these five produce a table that looks entirely reasonable and is wrong.",[576,577,578,579,580],"One price column. The page showed two numbers and the table kept one, so nobody can separate a permanent reprice from a three-day promotion, and the discount history is gone for good.","Interpreting availability at capture time. \"Ships in 2–3 weeks\" becomes true in an in_stock boolean, and six months later nobody can reconstruct what the page actually said.","Trusting JSON-LD without verifying it. It is the most stable source on the page and also the one most often left stale after a promotion — check the sale price against the rendered page on a handful of listings before you rely on the field.","Keeping the price as a string with its symbol attached. €1.299,00 and $1,299.00 both parse to 1299, and both parse to 1.299 if you guess the separator wrong. Keep the number and the currency in separate columns.","Reading a market you did not mean to read. Several large retailers decide currency and price from a cookie rather than from the URL, so a run can capture the wrong country's figure correctly, with no error raised anywhere.",{"type":127,"title":582,"intro":583,"items":584},"Check these before you scale up","Run one page per competitor and look at the output with your own eyes. Five minutes here saves a month.",[585,586,587,588,589],"A product that is on sale — do you get both prices, in the right columns?","A product that is out of stock — do you get a price at all, and is it the right one?","A multi-variant product — which variant did you capture, and did you mean to?","A product in a second currency or market domain, if the competitor has one","The most expensive product you track — decimal separator errors turn 1.299,00 into 1.29 and nothing downstream will question it",[591,592,593,594],"Never store one price column. List price and sale price separately; derive the discount later.","Capture availability as the exact words on the page and classify it downstream.","JSON-LD is the most stable source when it exists, but verify it — sale prices in particular are often wrong.","Selectors are cheap and brittle; a model reading the page is dearer and survives redesigns. Choose based on who is on call to fix it.",{"text":596,"label":597,"url":598},"You now have competitor rows. The next lesson is the one that decides whether any of them mean anything.","Lesson 5: matching their listings to your catalogue","/learn/competitor-price-monitoring/match-listings-to-your-catalogue",{"slug":600,"nav_title":601,"title":602,"summary":603,"time":604,"needs_account":73,"seo":605,"blocks":609,"takeaways":732,"next_step":737},"match-listings-to-your-catalogue","Matching to your catalogue","Matching competitor listings to your own catalogue","The stage that decides whether your feed is intelligence or fiction, and the denominator trick that makes bad match rates look good.","13 min",{"title":606,"description":607,"keywords":608},"Product Matching: Linking Competitor Listings to Your SKUs","How product matching really works, why match rate is meaningless without its denominator, and the four tiers of evidence from EAN down to a model's guess.","product matching, match competitor products, ean matching, product data matching, match rate, sku matching price monitoring",[610,614,622,625,655,661,679,684,721,729],{"type":80,"paragraphs":611},[612,613],"This is the lesson that decides whether everything before it was worth doing. A competitor price is only meaningful attached to one of your products, and deciding that their listing and your SKU are the same sellable thing is harder than it sounds, permanently imperfect, and almost never discussed honestly in a sales demo.","It is also the stage where you should be most suspicious of numbers, including your own.",{"type":80,"title":615,"paragraphs":616},"Four tiers of evidence",[617,618,619,620,621],"Matches are not binary. They sit on a ladder of confidence, and a well-run pipeline knows which rung each row is on.","Tier one is a shared global identifier: EAN, GTIN, UPC, ISBN. If both sides publish one and they agree, you are done — this is the only tier that is genuinely certain, and it is why lesson four told you to grab an EAN whenever the page shows one.","Tier two is a manufacturer part number plus brand. Strong, but MPNs get typed by hand into product feeds, so expect whitespace, hyphens and case to differ, and expect a small number of outright errors on both sides.","Tier three is attribute matching: brand, model, capacity, colour, pack size. This is where most real matching happens, and where most real errors live. Two litres against one litre, a twelve-pack against a single, last year's model number against this year's.","Tier four is fuzzy title similarity, usually with an embedding model. Useful for generating candidates a human then confirms. Not useful as a final answer, because it is confidently wrong in exactly the cases that cost the most money.",{"type":116,"variant":264,"title":623,"text":624},"The denominator question","\"We achieve a 92% match rate\" is not a claim until you ask: 92% of what? Of your whole catalogue, or of the products the competitor actually stocks, or of the products where both sides published an EAN? Those three denominators can differ by a factor of four on the same data. When a vendor quotes a match rate, ask what the bottom of the fraction is — and ask yourself the same question about your own reporting.",{"type":85,"title":626,"intro":627,"headers":628,"rows":631},"What a realistic match rate looks like","Measured against the honest denominator: products the competitor actually carries. These are the bands we see; your category decides where you land.",[629,630,238],"Category","Typical achievable",[632,636,640,644,648,652],[633,634,635],"Books, media, games","95%+","Universal identifiers, published by everyone, rarely wrong",[637,638,639],"Branded electronics","85-95%","MPNs are usually present and usually correct",[641,642,643],"Grocery and FMCG","70-90%","EANs are common, but pack size and multipacks create genuine ambiguity",[645,646,647],"Fashion","40-70%","Season codes, colour names and exclusive colourways. Hard, and honestly hard.",[649,650,651],"Bikes, furniture, configurables","30-60%","Variants and bundles mean \"the same product\" often is not a well-defined idea",[653,293,654],"Own-brand","There is nothing to match to. Exclude at scope time, per lesson two.",{"type":80,"title":656,"paragraphs":657},"Fix your own data first",[658,659,660],"The instinct is to blame the competitor's messy listings. In our experience the larger share of the problem is on the home side, and it is cheaper to fix.","Missing EANs in your own catalogue are the biggest single cause of poor matching, and most retailers have more holes than they think — the field exists, it is populated for 60% of lines, and nobody has looked. Pack sizes buried inside a free-text title rather than held as a number are the second. Duplicate internal SKUs for the same physical product are the third: the competitor has one listing, you have three, and your match rate is immediately capped at a third for those lines.","Before investing in better matching, spend a day measuring: what proportion of your in-scope SKUs have a populated, valid EAN? That number predicts your ceiling better than any vendor's algorithm.",{"type":210,"title":662,"items":663},"A workable process",[664,667,670,673,676],{"title":665,"text":666},"Match on identifiers first, and stop there where you can","Run the EAN and MPN passes, mark those rows as high confidence, and take them out of the pool. Depending on category this is anywhere from a third to nearly all of your matches, and it needs no review.",{"title":668,"text":669},"Generate candidates for the rest","Use title and attribute similarity to propose, say, the three most likely counterparts for each unmatched product. The goal here is recall, not precision — you want the right answer to be in the list, not to be alone in it.",{"title":671,"text":672},"Have a human confirm the candidates","Someone who knows the category can confirm or reject a proposed match in a couple of seconds from a title, an image and a price. A few hundred of those is an afternoon, it is a one-off cost per product, and the decisions are reusable forever.",{"title":674,"text":675},"Store the confirmed match, not the algorithm","Once a human has said \"their URL X is our SKU Y\", that is a fact you own. Persist it. Re-deriving it on every run is how a pipeline that worked last month starts producing different answers this month for no visible reason.",{"title":677,"text":678},"Review the unmatched pile monthly","Unmatched is not a failure state, it is a queue. It also carries a signal worth reading: a sudden growth in unmatched rows for one competitor usually means they have relabelled a range or changed their title format, not that your matcher broke.",{"type":80,"title":680,"paragraphs":681},"Decide what to do with a partial match",[682,683],"The genuinely difficult cases are not the failures, they are the near-misses. Their listing is the same model in a different colour. Theirs is a bundle with a case included. Theirs is the 2024 model and you stock the 2025.","There is no universally correct answer, but there is a correct process: decide the policy once, in writing, with the person who owns pricing — and record the reason a row was matched or excluded, so that six months later somebody can see why the 2024 model was being compared to the 2025 one. An undocumented matching policy becomes folklore within a quarter.",{"type":85,"title":685,"intro":686,"headers":687,"rows":690},"Worked example: the same match rate told three ways","A supplier reports 90%. The buyer counts 58.5%. Neither is lying; they are dividing by different numbers. Take 1,200 of your SKUs run against one competitor. The third figure is the one that belongs in a business case, because it is the only one you can reprice from without a person looking first.",[145,688,689],"Count","What it gives you",[691,694,698,702,706,710,713,717],[692,151,693],"SKUs you sent","the question you asked",[695,696,697],"Not stocked by this competitor at all","420","780 are even eligible to match",[699,700,701],"Eligible, but could not be resolved to a listing","78","702 matched",[703,704,705],"Of the 702, matched on a shared EAN","430","safe to act on unreviewed",[707,708,709],"Of the 702, matched on title and attributes only","272","needs a human before it moves a price",[711,287,712],"Vendor's headline: 702 ÷ 780","true, and not the number you need",[714,715,716],"Yours: 702 ÷ 1,200","58.5%","coverage of the question you asked",[718,719,720],"Confident coverage: 430 ÷ 1,200","about 36%","what you can actually price from",{"type":127,"title":177,"intro":722,"items":723},"Matching is where a feed stops being data and starts being fiction, and it does so without raising anything.",[724,725,726,727,728],"Accepting a match rate with no denominator. Ask \"out of what?\" before the figure is written down anywhere, because once it is in a slide it will be quoted for a year.","Letting tier-three matches move prices. Title similarity puts a 2-litre bottle next to a 1-litre bottle and a 2026 model next to a 2024 one, and both look exactly like a competitor undercutting you by half.","Re-deriving matches on every run. A human confirmed that pair once; if the pipeline rediscovers it nightly it will eventually rediscover it differently, and nobody will be able to say when the number changed.","Fixing the matcher instead of your own data. Missing EANs and inconsistent pack sizes in your own catalogue are usually the largest single cause, and they are the only part of this you fully control.","Counting a variant match as a product match. Their listing is the 128 GB model and yours is the 256 GB: same brand, near-identical title, different thing, and the report says you are 180 euros overpriced.",{"type":116,"variant":117,"title":730,"text":731},"Matching is bought separately, usually","Most collection vendors, including us by default, hand you the competitor's rows as the competitor presents them. Tying those rows to your catalogue is a different job with a different cost shape. If a tool implies both are included, find out which denominator they are quoting before you plan around it.",[733,734,735,736],"Matches sit on four tiers of evidence. Only a shared EAN or GTIN is certain.","A match rate without its denominator is not information. Ask what the bottom of the fraction is, every time.","The biggest fixable cause of poor matching is usually your own missing EANs and pack sizes.","Persist human-confirmed matches as facts. Never re-derive them on every run.",{"text":738,"label":739,"url":740},"A matched feed that silently loses a third of its rows is worse than no feed. Next: how to notice.","Lesson 6: scheduling and data quality","/learn/competitor-price-monitoring/schedule-runs-and-catch-silent-failure",{"slug":742,"nav_title":743,"title":744,"summary":745,"time":199,"needs_account":73,"seo":746,"blocks":750,"takeaways":850,"next_step":855},"schedule-runs-and-catch-silent-failure","Scheduling and data quality","Scheduling runs and catching silent data loss","How often to actually check, and the four alerts that catch a degrading feed before someone reprices from it.",{"title":747,"description":748,"keywords":749},"Price Monitoring Data Quality: Catching Silent Feed Failure","How often to run price checks, and the four alerts that catch a feed degrading quietly: row count, null rate, price jumps and staleness.","price monitoring data quality, scraper monitoring, silent data loss, scraping alerts, price feed reliability, how often to check competitor prices",[751,755,761,764,788,793,802,805,836,845],{"type":80,"paragraphs":752},[753,754],"There are two questions in this lesson and the second one matters more. How often should you check? And how will you find out when checking stops working?","Almost every team gets the first right by accident and the second wrong for months.",{"type":80,"title":756,"paragraphs":757},"How often",[758,759,760],"Match the cadence to the decision, not to what is technically possible. If your prices are reviewed weekly by a human, a daily feed is already giving them more than they can use, and an hourly feed is giving them seven times the bill for the same decision.","Daily, overnight, is right for the large majority of retail. Weekly is right for slow categories and for the long-tail competitors you decided in lesson two to watch rather than track. Intra-day genuinely matters in two situations: marketplace buy-box repricing, where the feedback loop is minutes, and flash-sale-heavy categories where a competitor's promotion starts and ends inside a day.","One detail worth getting right: run at the same time every day, and record the timestamp with the row. Prices change through the day, and a feed that samples at 03:00 on Monday and 14:00 on Tuesday produces price movements that are artefacts of your own schedule.",{"type":116,"variant":264,"title":762,"text":763},"The failure mode that matters","Scrapers rarely stop. They degrade. The run completes, the job is green, the file lands at the usual time — and it has 6,200 rows instead of 8,900 because a redesign broke one competitor's price element. Nothing alerts, because nothing errored. Three weeks later someone notices the feed has been missing a competitor's entire electronics range, and every repricing decision made in that window was taken on partial data.",{"type":85,"title":765,"intro":766,"headers":767,"rows":771},"Four alerts that catch almost everything","None of these are sophisticated. All of them are better than a green tick on a completed job.",[768,769,770],"Alert","Trigger","What it usually means",[772,776,780,784],[773,774,775],"Row count drop","Today's rows below 90% of the trailing 7-day median","A competitor redesigned, or added a bot wall, or you lost a category",[777,778,779],"Null rate spike","A field that is normally 98% populated falls below 90%","That specific field moved on the page; the rest of the row is fine",[781,782,783],"Implausible price move","A price changes by more than 50% overnight","Usually a decimal separator or a currency flip, occasionally a real clearance",[785,786,787],"Staleness","A tracked product has not returned a price for 3 consecutive runs","The URL is dead, or it is genuinely discontinued — both need action",{"type":80,"title":789,"paragraphs":790},"Alert per competitor, not per feed",[791,792],"This is the detail that makes the difference between alerts that work and alerts that get muted. If you have eight competitors and one of them breaks completely, your total row count drops by maybe twelve percent — comfortably inside the noise of a whole-feed threshold, and completely invisible.","Compute the same four checks per competitor, and a total failure becomes a hundred percent drop on one segment, which no threshold can miss. The cost is a slightly longer alert configuration; the benefit is catching the exact failure that a feed-level alert is structurally blind to.",{"type":127,"title":794,"items":795},"Things that will break your run, roughly in order of likelihood",[796,797,798,799,800,801],"A competitor redesigns their product template, usually without warning and usually in Q1","A bot wall appears, often seasonally, often only for datacentre IP ranges — meaning it works from your laptop and fails in production","A product URL starts redirecting to a category page, which returns a valid 200 and no price","A market or language switch changes the currency served from the same URL","Your own in-scope list grows and nobody updated the expected row count, so the baseline is wrong","A promotion changes the price element entirely for the duration of a sale, then changes back",{"type":116,"variant":117,"title":803,"text":804},"Test from where it runs","The second item on that list deserves its own warning. Bot walls very often key on the IP range, so a competitor's page that loads perfectly in your browser can return a 403 from a cloud server. Debugging that from your laptop produces a confident, wrong conclusion. Always reproduce the failure from the machine that actually makes the request.",{"type":85,"title":806,"intro":807,"headers":808,"rows":813},"Worked example: the morning a feed lost a quarter of itself","Alerts only work if someone wrote down what normal looks like first. One competitor over five days, with the row-count threshold set at 10% below the trailing median. No job failed in that week — every one of the five runs exited zero. Thursday's 531 was a category page that had started paginating differently, and without the row-count check the first sign of it would have been a buyer asking why a product had vanished from the sheet.",[809,810,811,812],"Day","Rows returned","Null price rate","Verdict",[814,819,823,827,832],[815,816,817,818],"Monday","712","1.1%","normal",[820,821,822,818],"Tuesday","709","0.8%",[824,825,826,818],"Wednesday","714","1.0%",[828,829,830,831],"Thursday","531","1.2%","alert — 25% below the 712 median",[833,834,826,835],"Friday","528","second day running — now an incident",{"type":127,"title":177,"intro":837,"items":838},"Every item here describes a pipeline whose dashboard is entirely green.",[839,840,841,842,843,844],"Alerting on the whole feed instead of per competitor. One competitor out of five dropping to zero is a 20% dip in the total — under almost any sensible threshold, and therefore invisible.","Treating a green job as a good run. Exit code zero means the code finished, not that the page still contained what it used to contain.","Testing from your laptop. The competitor serves your office address happily and blocks the data centre the job runs in, so the check passes and the run does not.","Setting thresholds against a fixed number rather than a trailing median. Catalogues grow, so a static floor either fires every week or stops firing altogether.","Deleting failed runs. A table containing only successes cannot tell you whether a product was out of stock on Thursday or whether Thursday never happened.","Setting the price-jump threshold too tight. At 5% every promotional weekend is an incident and the alert gets muted; at 60% you still catch the decimal-point errors and currency mix-ups, which are the ones that actually corrupt a reprice.",{"type":80,"title":846,"paragraphs":847},"Keep the history, including the gaps",[848,849],"Store every observation with its date, rather than overwriting a current price column. The overwrite feels tidier and destroys the asset: price history is where you see that a rival's discount is cyclical, that they follow you within two days, that a category has been drifting down for a quarter.","Store the misses too. A row that returned no price on a given day is information — it may be a stock-out, a delisting or a broken scraper — and a table that only contains successes cannot tell you which. A gap you cannot see is a gap you will interpret as stability.",[851,852,853,854],"Match frequency to the decision cycle. Daily overnight suits most retail; hourly is for buy-box repricing.","Scrapers degrade far more often than they stop. A green job is not evidence of a good feed.","Run row-count, null-rate, price-jump and staleness checks per competitor, not across the whole feed.","Keep dated history including the failures. A table of successes cannot distinguish a stock-out from a broken run.",{"text":856,"label":857,"url":858},"A trustworthy feed in a system nobody opens still changes nothing. Next: getting it where the decision happens.","Lesson 7: exporting to Sheets, BI and your ERP","/learn/competitor-price-monitoring/export-to-sheets-bi-and-erp",{"slug":860,"nav_title":861,"title":862,"summary":863,"time":72,"needs_account":338,"seo":864,"blocks":868,"takeaways":965,"next_step":970},"export-to-sheets-bi-and-erp","Getting the data out","Getting the data into Sheets, BI or your ERP","Four delivery routes ranked by how likely they are to actually get used, and the column contract that stops downstream jobs breaking.",{"title":865,"description":866,"keywords":867},"Exporting Competitor Price Data to Sheets, BI or ERP","Four ways to deliver a price feed — spreadsheet, CSV drop, REST API, warehouse — and the stable column contract that keeps downstream jobs from breaking.","export price data, price feed api, competitor prices google sheets, price data bi dashboard, csv price export, rest api price monitoring",[869,873,896,901,906,915,951,960],{"type":80,"paragraphs":870},[871,872],"A correct, well-matched, well-monitored price feed that lands somewhere nobody looks has the same business value as no feed at all. This lesson is short because the principle is simple: deliver into the tool where the pricing decision is already being made, not into a new tool you are hoping people will adopt.","In practice that is almost always a spreadsheet, and there is no shame in it.",{"type":85,"title":874,"headers":875,"rows":879},"Four routes",[876,877,878],"Route","Good for","The catch",[880,884,888,892],[881,882,883],"Spreadsheet","A pricing manager who already works in one","Breaks at a few hundred thousand rows; versions multiply",[885,886,887],"Scheduled CSV or Excel drop","Feeding an existing internal job","Somebody has to own the filename convention and the failure case",[889,890,891],"REST API","Your own application or an automation tool","Requires a developer once; then it is the lowest-maintenance option",[893,894,895],"Warehouse table","Teams already running BI","Highest setup cost, and only pays off if the dashboards get opened",{"type":80,"title":897,"paragraphs":898},"The column contract",[899,900],"Whatever the route, the single most valuable property of a feed is that its columns do not move. Downstream jobs — a spreadsheet formula, an ETL step, a repricing script — are written against column names and positions, and they break silently when those change. A column that disappears usually produces a blank rather than an error, and a blank price reads as free.","So pin the contract. Every run returns the same columns, in the same order, with the same names, whether or not a particular page had a value for them. An empty cell is a valid answer and must be distinguishable from a missing column. Adding a new column at the end is safe; renaming or reordering is not, and should be treated as a versioned change with a deprecation window, however informal.",{"type":116,"variant":364,"title":530,"text":902,"cta":903},"Every run is retrievable over REST with a stable column set, or downloadable as CSV or Excel from the portal. The columns come from the schema you defined in lesson four, which is what keeps them identical across every competitor in the group — the thing that makes a single downstream job work for all of them rather than one per site.",{"label":904,"url":905},"See how the exports work","/how-it-works",{"type":127,"title":907,"intro":908,"items":909},"Fields the consumer of the feed will ask for within a week","Include them from the start; retrofitting them means reprocessing history.",[910,911,912,913,914],"The observation timestamp, not just the date — so a reader can tell a morning sample from an evening one","The competitor name as a clean key, not as a URL to be parsed","Your own SKU alongside their identifier, so no downstream join is needed","The match tier, so a reader can see which rows are EAN-certain and which are a judgement call","The source URL, so a surprising number can be checked by clicking it",{"type":85,"title":916,"intro":917,"headers":918,"rows":921},"Worked example: a column contract in full","The whole contract for a morning sheet — every column, in a fixed order, emitted on every run whether or not there is a value for it. The right-hand column is the part that matters: it is what everything downstream is entitled to assume. Guarantees five and six exist because a spreadsheet treats an empty cell as zero in most arithmetic, and a zero price reads as free.",[919,540,920],"#","Guarantee",[922,925,929,933,937,939,942,945,948],[923,567,924],"1","Always present, UTC, ISO 8601. Never the time the sheet was opened.",[926,927,928],"2","competitor_key","A stable slug, never the display name. Renaming a shop must not break a formula.",[930,931,932],"3","your_sku","Your identifier, not theirs. This is the join key for everything downstream.",[934,935,936],"4","match_tier","1 to 4, present even when the match is certain.",[155,481,938],"Empty if the page showed no struck-through price. Empty, not zero.",[940,548,941],"6","Empty if there was no promotion. Empty, not a copy of list_price.",[943,484,944],"7","ISO code, on every row, including rows from your home market.",[946,555,947],"8","Verbatim from the page, uninterpreted.",[949,571,950],"9","The page the row came from, so an argument can be settled in one click.",{"type":127,"title":177,"intro":952,"items":953},"Delivery is the stage where a technically correct feed becomes a feed nobody uses.",[954,955,956,957,958,959],"Emitting a column only when it has data. The header row shifts, every formula in the sheet moves one column left, and the error stays silent for a week.","Building a dashboard first. It gets opened on launch day and at the quarterly review, while the pricing decision carries on happening in a spreadsheet nobody told you about.","Delivering the competitor's identifier and not yours. Every consumer then has to redo the join, and each of them will do it slightly differently.","Using zero as the empty value. A zero price is a 100% undercut, and most repricing rules will act on it before anyone notices.","Overwriting yesterday's file. The first time somebody asks \"when did they drop it?\", the only honest answer is \"sometime in the last month\".","Emailing the whole catalogue every morning. It stops being opened in week three. Send the rows that moved and link to the full sheet.",{"type":80,"title":961,"paragraphs":962},"Ship one view, not a platform",[963,964],"The strongest first deliverable we have seen is a single sheet, refreshed every morning, with one row per in-scope product and one column per competitor, cells coloured by whether you are above or below. No navigation, no filters, no login. It is unglamorous and people open it.","The elaborate version — the dashboard with the drill-downs and the trend charts — is the right second deliverable, once the daily sheet has proven somebody cares. Built first, it is usually the thing that makes the project look expensive and optional at the same time.",[966,967,968,969],"Deliver into the tool where the decision already happens. That is usually a spreadsheet.","Freeze the column contract: same names, same order, every run. A missing column reads as a blank, and a blank price reads as free.","Include timestamp, competitor key, your SKU, match tier and source URL from day one.","Ship one morning sheet before building a dashboard. Adoption first, sophistication second.",{"text":971,"label":972,"url":973},"Last lesson: turning a column of competitor prices into a decision that does not destroy your margin.","Lesson 8: from price data to repricing rules","/learn/competitor-price-monitoring/turn-price-data-into-repricing-rules",{"slug":975,"nav_title":976,"title":977,"summary":978,"time":337,"needs_account":73,"seo":979,"blocks":983,"takeaways":1078,"next_step":1083},"turn-price-data-into-repricing-rules","From data to decisions","Turning price data into repricing decisions","Why \"match the cheapest\" destroys margin, what a rule needs besides a competitor price, and how to start without automating anything.",{"title":980,"description":981,"keywords":982},"Turning Competitor Price Data Into Repricing Rules","Why matching the cheapest competitor destroys margin, the guard conditions every repricing rule needs, and how to start with recommendations instead of automation.","repricing rules, competitor based pricing, dynamic pricing rules, price matching strategy, repricing strategy ecommerce, margin floor pricing",[984,988,995,1014,1017,1036,1042,1066,1075],{"type":80,"paragraphs":985},[986,987],"You have a trustworthy daily feed of matched competitor prices landing where people can see it. The last question is what anyone should do with it, and this is where a good data project can still destroy value.","The naive rule — be one cent below the cheapest competitor — is the one everybody writes first. It is worth understanding precisely why it is bad before replacing it.",{"type":80,"title":989,"paragraphs":990},"Why matching the cheapest fails",[991,992,993,994],"It hands your pricing to your least informed competitor. Somewhere in any set of rivals is a seller clearing stock, miscalculating shipping, or simply making a mistake. A rule that follows the minimum follows that seller, and the rest of the market follows you.","It ignores everything a shopper is actually comparing. Delivery speed and cost, returns, stock availability, whether the buyer trusts the site. Being three percent dearer with next-day delivery and real stock is a winning position in most categories; a rule that only reads a number cannot see that.","It races. If two sellers in a market both run undercut rules, the price falls to the floor within days, and neither of them chose that outcome. This is visible and documented on marketplaces, and it is the main reason serious repricing systems include a floor and a cooldown.","And it reacts to noise. A competitor who is out of stock still displays a price, often a stale one. Following it means repricing against a product nobody can buy.",{"type":210,"title":996,"intro":997,"items":998},"What a usable rule contains","A competitor price is one input. These are the others.",[999,1002,1005,1008,1011],{"title":1000,"text":1001},"A margin floor, per product","The price below which you would rather lose the sale. Non-negotiable, checked last, and it should override every other part of the rule. Most repricing disasters are a missing or a global-rather-than-per-product floor.",{"title":1003,"text":1004},"A reference set, not a minimum","Decide which competitors the rule may react to — usually the two or three from lesson two that genuinely take your customers — and use a position within that set, such as \"the second cheapest\" or \"the median\", rather than the absolute minimum. It is dramatically more stable.",{"title":1006,"text":1007},"An availability condition","Ignore the price of any competitor whose page says out of stock. This single condition removes a large share of spurious triggers, and it is only possible because lesson four told you to keep the availability text.",{"title":1009,"text":1010},"A movement threshold and a cooldown","Do not act on a change smaller than, say, one percent, and do not change the same product's price more than once in a defined window. Thresholds stop you chasing rounding; cooldowns stop two automated systems from spiralling.",{"title":1012,"text":1013},"A total-cost comparison where you can get it","Where you captured shipping, compare delivered price rather than shelf price. A rival who is two percent cheaper and charges for delivery is not cheaper, and a rule reading only the shelf price will tell you they are.",{"type":116,"variant":117,"title":1015,"text":1016},"Start with recommendations, not automation","Run the rule in advisory mode for a few weeks: it produces a list of suggested price changes each morning and a human approves them. You will find out quickly which of your rules are wrong, and you will find out on a screen rather than on your P&L. The approval list also tells you something useful — if a person approves 98% of suggestions without reading them, the rule is ready to automate; if they override a third of them, the rule is missing a condition they are applying from memory.",{"type":85,"title":1018,"intro":1019,"headers":1020,"rows":1023},"Three rules that tend to hold up","Illustrative shapes rather than recommendations. The right rule depends on your category and your position in it.",[381,1021,1022],"Rule shape","Why it works",[1024,1028,1032],[1025,1026,1027],"You are the service leader","Stay within +4% of the reference median, never below the floor","Captures the price-sensitive shopper without giving away the premium you have earned",[1029,1030,1031],"Clearing end-of-life stock","Match the second cheapest in-stock reference, floor at cost","Moves units without leading the market down on live lines",[1033,1034,1035],"Enforcing a supplier's minimum price","Never price below MAP; alert when a reference breaks it","The output is an evidence trail for your supplier, not a price change",{"type":80,"title":1037,"paragraphs":1038},"Measure the rule, not the feed",[1039,1040,1041],"The last trap is reporting on the wrong thing. It is tempting to report pipeline health — rows collected, match rate, uptime — because those numbers are easy and they go up. They are not the point.","The questions worth answering quarterly are: how many of our in-scope products are currently priced outside the band we intended, how quickly do we react when a reference competitor moves, and what happened to margin and volume on the lines the rule touched compared with the lines it did not. That last comparison is the only honest measure of whether any of this worked, and it requires that you deliberately leave a set of products out of the rule so there is something to compare against.","Hold one back. A control group of a few hundred SKUs costs you almost nothing and is the difference between knowing the project paid for itself and believing it.",{"type":85,"title":1043,"intro":1044,"headers":1045,"rows":1050},"Worked example: one SKU through three rules","A product you sell at €99.00 that costs you €72.00, so €27.00 of margin, with a floor set at 12% — €80.64. Three competitors are visible this morning: €104.00 in stock, €96.50 in stock, and €84.00 marked \"2–3 weeks\". The first rule cuts your margin by more than half to beat a price the shopper cannot buy today. The third rule raises the price.",[1046,1047,1048,1049],"Rule","Reference price","New price","Margin kept",[1051,1056,1061],[1052,1053,1054,1055],"Match the cheapest, any availability","€84.00","€83.90","€11.90",[1057,1058,1059,1060],"Match the cheapest that is in stock","€96.50","€96.40","€24.40",[1062,1063,1064,1065],"Sit 1% under the median of in-stock rivals","median of €104.00 and €96.50 = €100.25","€99.25","€27.25",{"type":127,"title":177,"intro":1067,"items":1068},"Each of these is a rule that works on the day it is written and costs money a month later.",[1069,1070,1071,1072,1073,1074],"No availability condition. The cheapest figure on the page is very often the price of something nobody can ship, and it drags your whole range down behind it.","No floor. A rule without a per-product margin floor will follow a competitor clearing stock all the way into a loss, at machine speed, overnight.","No movement threshold. Reacting to a four-cent change generates hundreds of price updates a day, each one a row in a channel feed and a fresh chance to be rejected.","No cooldown. Two automated repricers pointed at each other will walk a price downwards in a loop, and the loop finishes long before anyone reads the alert.","Automating in week one. Run it advisory and count the overrides: disagreement on one row in twenty means it is ready, and one in four means the rule is wrong, not the human.","No control group. Without a held-out set of SKUs the rule never has to prove it beat doing nothing, and nobody can answer whether the pipeline paid for itself.",{"type":116,"variant":364,"title":1076,"text":1077},"Where we fit","Scrapewise covers stage two of lesson one: collecting the competitor rows, daily, from any public site, with a stable column set. Matching and repricing are yours — or a specialist tool's. We have written this course the way we have because a customer who understands the other three stages gets more out of the one we do.",[1079,1080,1081,1082],"Matching the cheapest competitor hands your pricing to whoever is clearing stock or making a mistake.","A usable rule needs a per-product margin floor, a chosen reference set, an availability condition, a movement threshold and a cooldown.","Run it advisory for a few weeks. The override rate tells you whether it is ready to automate.","Keep a control group. It is the only way to know whether the whole pipeline paid for itself.",{"text":1084,"label":1085,"url":534},"That is the course. If you want the collection stage handled, the retailer write-ups show what a run on a specific site returns and what it costs.","Browse the retailer write-ups",[1087,1098,1139,1178,1217,1255],{"order":5,"slug":4,"title":17,"subtitle":18,"cardText":8,"level":6,"time":7,"lessonCount":1088,"lessons":1089},8,[1090,1091,1092,1093,1094,1095,1096,1097],{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":195,"navTitle":196,"title":197,"summary":198,"time":199,"needsAccount":73},{"slug":333,"navTitle":334,"title":335,"summary":336,"time":337,"needsAccount":338},{"slug":450,"navTitle":451,"title":452,"summary":453,"time":337,"needsAccount":338},{"slug":600,"navTitle":601,"title":602,"summary":603,"time":604,"needsAccount":73},{"slug":742,"navTitle":743,"title":744,"summary":745,"time":199,"needsAccount":73},{"slug":860,"navTitle":861,"title":862,"summary":863,"time":72,"needsAccount":338},{"slug":975,"navTitle":976,"title":977,"summary":978,"time":337,"needsAccount":73},{"order":1099,"slug":1100,"title":1101,"subtitle":1102,"cardText":1103,"level":1104,"time":1105,"lessonCount":1106,"lessons":1107},2,"ai-agent-web-data-mcp","Give your AI agent live web data via MCP","Ask an assistant what a product costs today and you will usually get a number. It is often wrong, and it is always wrong in the same way: the model is reconstructing a plausible price from training data rather than looking at a page. This course is about closing that gap properly — what the Model Context Protocol actually is, how to wire a server into a client, how to design tools a model can use without hand-holding, and what to put in place before an agent spends your money.","Your agent is confidently wrong about prices because it has never seen one. What MCP is, how to connect a server, how to design tools a model can actually use, and the guardrails you need before you let it loose.","Comfortable editing a config file","6 lessons, about 60 minutes",6,[1108,1113,1119,1124,1129,1134],{"slug":1109,"navTitle":1110,"title":1111,"summary":1112,"time":72,"needsAccount":73},"what-is-mcp","What MCP is","What MCP actually is, in plain terms","The Model Context Protocol described without jargon: what problem it solves, its three primitives, and when it is the wrong tool.",{"slug":1114,"navTitle":1115,"title":1116,"summary":1117,"time":1118,"needsAccount":73},"why-agents-get-live-data-wrong","Why agents get it wrong","Why your agent's answer about a price is wrong","Four distinct failure modes that all look identical from the outside, and how to tell which one you have before you try to fix it.","10 min",{"slug":1120,"navTitle":1121,"title":1122,"summary":1123,"time":1118,"needsAccount":73},"connect-an-mcp-server","Connecting a server","Connecting an MCP server and proving it works","The config for local and remote servers, the four things that go wrong, and how to verify the tools registered rather than assuming.",{"slug":1125,"navTitle":1126,"title":1127,"summary":1128,"time":199,"needsAccount":73},"design-tools-an-agent-can-use","Designing usable tools","Designing tools an agent can actually use","A connected server is not a useful server. The model only sees your tool names, descriptions and parameter schemas, so those three things are the entire user interface. Here is what makes a tool get called correctly and what makes it get ignored.",{"slug":1130,"navTitle":1131,"title":1132,"summary":1133,"time":337,"needsAccount":338},"give-an-agent-a-scraper","Giving an agent a scraper","Giving an agent a real price feed","A worked example. Connect the ScrapeWise MCP server to a client, let the agent read a live scraper's output, and watch where the hand-off between \"the data is right\" and \"the answer is right\" actually breaks.",{"slug":1135,"navTitle":1136,"title":1137,"summary":1138,"time":199,"needsAccount":73},"guardrails-cost-and-untrusted-content","Guardrails and cost","Guardrails, cost control and untrusted content","Live web access turns an agent into something that can spend money and read text written by strangers. Neither is a reason not to do it. Both are reasons to put limits in before you need them.",{"order":1140,"slug":1141,"title":1142,"subtitle":1143,"cardText":1144,"level":1145,"time":1146,"lessonCount":1106,"lessons":1147},3,"product-data-api","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.","Comfortable with HTTP and JSON","6 lessons, about 70 minutes",[1148,1153,1158,1163,1168,1173],{"slug":1149,"navTitle":1150,"title":1151,"summary":1152,"time":1118,"needsAccount":73},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.",{"slug":1154,"navTitle":1155,"title":1156,"summary":1157,"time":1118,"needsAccount":73},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":1159,"navTitle":1160,"title":1161,"summary":1162,"time":337,"needsAccount":338},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.",{"slug":1164,"navTitle":1165,"title":1166,"summary":1167,"time":199,"needsAccount":73},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.",{"slug":1169,"navTitle":1170,"title":1171,"summary":1172,"time":199,"needsAccount":73},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"slug":1174,"navTitle":1175,"title":1176,"summary":1177,"time":337,"needsAccount":338},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",{"order":1179,"slug":1180,"title":1181,"subtitle":1182,"cardText":1183,"level":1184,"time":1185,"lessonCount":1106,"lessons":1186},4,"keep-scrapers-alive","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.","You already have something running","6 lessons, about 65 minutes",[1187,1192,1197,1202,1207,1212],{"slug":1188,"navTitle":1189,"title":1190,"summary":1191,"time":1118,"needsAccount":73},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.",{"slug":1193,"navTitle":1194,"title":1195,"summary":1196,"time":337,"needsAccount":73},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.",{"slug":1198,"navTitle":1199,"title":1200,"summary":1201,"time":199,"needsAccount":73},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.",{"slug":1203,"navTitle":1204,"title":1205,"summary":1206,"time":199,"needsAccount":73},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":1208,"navTitle":1209,"title":1210,"summary":1211,"time":199,"needsAccount":73},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":1213,"navTitle":1214,"title":1215,"summary":1216,"time":1118,"needsAccount":73},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"order":1218,"slug":1219,"title":1220,"subtitle":1221,"cardText":1222,"level":1223,"time":1146,"lessonCount":1106,"lessons":1224},5,"matching-products-across-sites","Match the same product across different sites","A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.","Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.","You have data from more than one site",[1225,1230,1235,1240,1245,1250],{"slug":1226,"navTitle":1227,"title":1228,"summary":1229,"time":1118,"needsAccount":73},"why-matching-is-the-hard-part","Why matching is hard","Why matching is the hard part","The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.",{"slug":1231,"navTitle":1232,"title":1233,"summary":1234,"time":199,"needsAccount":73},"identifiers-first-gtin-ean-mpn","Identifiers first","Identifiers first: GTIN, EAN, UPC and MPN","What each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.",{"slug":1236,"navTitle":1237,"title":1238,"summary":1239,"time":337,"needsAccount":73},"when-there-is-no-barcode","No barcode","Matching when there is no barcode","Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.",{"slug":1241,"navTitle":1242,"title":1243,"summary":1244,"time":337,"needsAccount":73},"variants-bundles-and-multipacks","Variants and packs","Variants, bundles and multipacks","The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.",{"slug":1246,"navTitle":1247,"title":1248,"summary":1249,"time":199,"needsAccount":338},"score-confidence-and-build-a-review-queue","Confidence and review","Confidence scores and a review queue worth using","Why one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.",{"slug":1251,"navTitle":1252,"title":1253,"summary":1254,"time":199,"needsAccount":73},"measure-your-match-rate-honestly","Measure it honestly","Measure your match rate honestly","The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.",{"order":1106,"slug":1256,"title":1257,"subtitle":1258,"cardText":1259,"level":1260,"time":1261,"lessonCount":1218,"lessons":1262},"web-scraping-legal-and-ethical","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.","No legal background assumed","5 lessons, about 55 minutes",[1263,1268,1273,1278,1283],{"slug":1264,"navTitle":1265,"title":1266,"summary":1267,"time":1118,"needsAccount":73},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.",{"slug":1269,"navTitle":1270,"title":1271,"summary":1272,"time":199,"needsAccount":73},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.",{"slug":1274,"navTitle":1275,"title":1276,"summary":1277,"time":199,"needsAccount":73},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"slug":1279,"navTitle":1280,"title":1281,"summary":1282,"time":199,"needsAccount":73},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.",{"slug":1284,"navTitle":1285,"title":1286,"summary":1287,"time":337,"needsAccount":73},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.",1791047866688]