[{"data":1,"prerenderedAt":992},["ShallowReactive",2],{"learn-course-product-data-api":3,"learn-courses":781},{"slug":4,"order":5,"level":6,"time":7,"card_text":8,"seo":9,"hero":15,"outcomes":25,"who":34,"syllabus":45,"faq":48,"lessons":65},"product-data-api",3,"Comfortable with HTTP and JSON","6 lessons, about 70 minutes","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.",{"title":10,"description":11,"keywords":12,"og_title":13,"og_description":14},"Product Data API: A Free 6-Lesson Course for Developers","How to pull product, price and stock data over an API. Authentication, declaring fields, asynchronous runs, pagination, retries without double billing, and putting the feed into your stack. Free, ungated.","product data api, price api, ecommerce api, web scraping api, api key authentication, api documentation, scraper api tutorial, rest api product data","A free developer course on pulling product data over an API","Six written lessons: when an API beats a scraper, auth, declaring fields, async runs, error handling, and shipping the feed into your stack.",{"badge":16,"title":17,"subtitle":18,"cta_primary":19,"cta_secondary":22},"Course three","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.",{"label":20,"url":21},"Start with lesson one","/learn/product-data-api/when-an-api-beats-a-scraper",{"label":23,"url":24},"See what a run returns","/custom-scrapers",{"title":26,"items":27},"What you will be able to do",[28,29,30,31,32,33],"Decide between an official retailer API, a web data API and writing your own scraper, with reasons you can defend in review","Make an authenticated request and read the response without guessing what the fields mean","Declare a schema so the extractor returns the fields you need rather than the ones it felt like returning","Work with asynchronous runs: start, poll, and handle a job that finishes half-done","Retry a failed request without being billed twice for the same page","Land the result in a warehouse table that does not quietly drift out of date",{"title":35,"for_title":36,"for":37,"not_title":41,"not_for":42},"Who this is for","Written for",[38,39,40],"Developers who have been told to \"get the competitor prices\" and have discovered the retailer has no public API","Data engineers deciding whether to own the collection layer or buy it","Backend teams wiring a third-party feed into an existing pipeline and wanting to know where it will break","Not written for",[43,44],"Readers looking for a no-code setup — course one covers the same ground without a terminal","Anyone wanting a specific vendor's endpoint reference. This is the shape of the problem; the reference lives in that vendor's docs.",{"title":46,"intro":47},"The six lessons","Lessons one and two are short and assume nothing but curl. From three onwards you will get more out of it with a key in your hand.",{"badge":49,"title":50,"description":51,"items":52},"FAQ","Before you start","The questions developers ask in the first ten minutes.",[53,56,59,62],{"title":54,"description":55},"Does this teach a specific vendor's API?","No, deliberately. Endpoint names change and a course that hardcodes them rots. What does not change is the shape: authenticate, declare what you want, start a run, poll it, handle partial results, retry safely. Learn the shape and any vendor's reference becomes a lookup rather than a tutorial.",{"title":57,"description":58},"Why is it not just a GET that returns the price?","Because fetching a page takes seconds and fetching ten thousand takes hours. Any API that hides that behind a synchronous request is either very slow or quietly returning you a cached number. Lesson four is about why the asynchronous shape exists and how to work with it instead of around it.",{"title":60,"description":61},"Do I need a Scrapewise account?","Lessons three and six include a Scrapewise-specific walkthrough and are labelled on the page. The other four are method and apply to whatever you are calling. A new account starts with five free requests and no card if you want to follow along.",{"title":63,"description":64},"Can I just scrape it myself in Python?","Often yes, and lesson one is honest about when that is the right call. The threshold is not technical skill — it is how many distinct sites you need and how much you mind being the person who gets paged when one of them redesigns. One site, one developer, no deadline: write it yourself.",[66,188,302,406,538,659],{"slug":67,"nav_title":68,"title":69,"summary":70,"time":71,"needs_account":72,"seo":73,"blocks":77,"takeaways":179,"next_step":184},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.","10 min",false,{"title":74,"description":75,"keywords":76},"Product API vs Web Scraper vs Dataset: How to Choose","Official retailer APIs, web data APIs and self-written scrapers compared on coverage, maintenance and cost. The question that actually decides it.","product api vs scraper, web data api, retailer api access, build or buy scraper, ecommerce data sourcing",[78,83,113,119,124,134,164,173],{"type":79,"paragraphs":80},"prose",[81,82],"If you search for \"Walmart API documentation\" or \"Target API key\" you will find thousands of other people who searched for the same thing. You will also find that the answer is usually some version of \"apply to be a partner\", and that the partner API, when you get it, covers your own listings rather than everybody else's. This is not an oversight. A retailer publishes an API so that suppliers can manage their own catalogue; publishing one that lets a competitor read every price would be a strange thing to do on purpose.","So the real choice is not between an official API and something else. It is between three options, and most teams pick one before they have articulated what they are actually optimising for.",{"type":84,"title":85,"intro":86,"headers":87,"rows":92},"table","The three ways to get product data","Coverage means which sites you can reach. Maintenance means who gets paged when one of them changes.",[88,89,90,91],"Route","Coverage","Maintenance","Best when",[93,98,103,108],[94,95,96,97],"Official retailer API","One retailer, usually only your own listings","Theirs, and it is versioned","You are a seller on that platform managing your own catalogue",[99,100,101,102],"Web data API","Any public page, if the vendor can reach it","Theirs, but you still own the schema","You need several sites and want one integration",[104,105,106,107],"Your own scraper","Whatever you build, exactly","Yours, forever","One or two sites, stable, and you want full control",[109,110,111,112],"A bought dataset","Broad but fixed, and already stale","Nobody's — it is a file","You need a one-off snapshot for analysis, not a feed",{"type":79,"title":114,"paragraphs":115},"The question that actually decides it",[116,117,118],"Not \"can we build this\". You can. A competent developer can read a price off a page in an afternoon, and the first version will work.","The question is: how many distinct sites, and how long does this have to keep working? Those two numbers together are the whole decision. One site for a quarter is a script. Forty sites indefinitely is an operations problem that happens to involve HTTP, and the cost is not the code — it is the on-call rotation, the day a redesign breaks nine of them at once, and the proxy bill you did not budget for.","The honest crossover, from running both: somewhere around three to five sites, or the moment the feed becomes something a person makes a pricing decision from. Below that, write it. Above that, the integration you maintain should be one API rather than forty parsers.",{"type":120,"variant":121,"title":122,"text":123},"callout","warning","The cost nobody puts in the build-versus-buy spreadsheet","Build estimates almost always cost the extraction and forget the detection. A scraper that silently returns four hundred rows instead of two thousand does not raise an exception — it returns successfully, with less. Somebody has to notice. Building that noticing is usually larger than building the scraper, and it is the part that gets cut when the deadline moves.",{"type":125,"title":126,"intro":127,"items":128},"list","Signs you have outgrown your own scraper","None of these is fatal on its own. Two or more and the maintenance has become the project.",[129,130,131,132,133],"You have a file called something like fixes.py and nobody remembers what half of it is working around","A redesign on one site has broken your run more than twice this year","You are maintaining a proxy pool, and somebody has had to think about residential versus datacenter IPs","The feed has been wrong in production and the first person to notice was not you","You have started writing per-site exceptions into what was supposed to be a generic parser",{"type":84,"title":135,"intro":136,"headers":137,"rows":141},"Worked example: twelve months of six sites, both ways","The build estimate that loses money is the one that prices the first version and then stops. Six competitor sites, a developer at a loaded cost of €400 a day, and a year of keeping it running. Look at which rows do the damage: the build is 6 days and the year is 20, and the other 14 all arrive after the project was marked done. At one site for one month, build — the maintenance tail never shows up. The tail is the entire decision.",[138,139,140],"","Build it yourself","Buy a web data API",[142,146,150,153,156,160],[143,144,145],"Initial extraction, six sites","6 days — €2,400","1 day of integration — €400",[147,148,149],"Blocking, retries, proxies, scheduling","8 days — €3,200","included",[151,152,149],"Breakages, assuming each site changes twice a year","12 × half a day — €2,400",[154,155,149],"Someone on call to notice a breakage","never in the estimate, and the real cost",[157,158,159],"Twelve-month labour","20 days — €8,000","1 day — €400",[161,162,163],"Infrastructure and fetches","proxies and servers on top of the above","the usage line you compare against",{"type":125,"title":165,"intro":166,"items":167},"What usually goes wrong","Build-versus-buy goes wrong in the estimate far more often than in the engineering.",[168,169,170,171,172],"Pricing the build and forgetting the year. The first working version is 6 of those 20 days; the remaining 14 arrive quietly, one afternoon at a time.","Assuming the retailer has an API. Most official retailer APIs exist to show you your own listings, which is the opposite of what price monitoring needs — a deliberate design decision rather than an oversight.","Counting developer days and ignoring the on-call. The expensive part of a breakage is the eleven days before anyone noticed, not the half day it took to fix.","Comparing a build quote against a list price with no volume attached. The vendor line is pages times frequency; without that multiplication you are comparing a number to a different kind of number.","Treating it as a technical decision. The question is how many sites, for how long, and who is on call when one changes — none of which are about whether you can write the parser.",{"type":79,"title":174,"paragraphs":175},"What a web data API is actually doing for you",[176,177,178],"Three things, and it is worth being precise because the pricing follows directly from them.","It fetches the page, which is the part that involves proxies, headers, retries and occasionally a headless browser. It extracts the fields, which is the part that involves knowing where a price lives on a page that has never been seen before. And it keeps doing both when the site changes, which is the part you are really paying for.","Everything else — scheduling, exports, dashboards — is convenience. If a vendor is strong on the dashboard and vague about what happens when a site redesigns, you are being sold the wrong half.",[180,181,182,183],"Official retailer APIs mostly expose your own listings, not your competitors'. That is by design, not an oversight.","The decision is sites multiplied by duration, not technical difficulty.","Build estimates price the extraction and forget the detection; detection is the expensive half.","What you pay a web data API for is the third thing: still working after the site changes.",{"text":185,"label":186,"url":187},"Next: the first authenticated call, and reading the response without guessing.","Lesson 2: authentication and your first call","/learn/product-data-api/authentication-and-your-first-call",{"slug":189,"nav_title":190,"title":191,"summary":192,"time":71,"needs_account":72,"seo":193,"blocks":197,"takeaways":293,"next_step":298},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"title":194,"description":195,"keywords":196},"API Key Authentication for Data APIs: Bearer Tokens Explained","How API key authentication works for product data APIs, bearer tokens versus query-string keys, key scoping and rotation, and reading your first response.","api key, bearer token, api authentication, api access, rest api key, api documentation, authorization header",[198,203,210,218,221,227,259,279,288],{"type":79,"paragraphs":199},[200,201,202],"Almost every data API authenticates one of two ways, and the difference matters more than it looks.","The common one is a bearer token in a header. You send Authorization: Bearer followed by the key, the server checks it, and the key never appears in the URL. The other is a key in the query string, which exists because it is trivially easy to test in a browser address bar and is therefore popular in quickstarts.","Prefer the header every time it is offered. A key in a query string ends up in server access logs, in browser history, in the Referer header of any outbound link, in your error-tracking tool's breadcrumb trail, and in the screenshot somebody pastes into a ticket. A key in a header ends up in none of those by default.",{"type":204,"title":205,"intro":206,"language":207,"code":208,"caption":209},"code","The shape of a first request","Nothing vendor-specific here. Replace the base URL and the path with whatever your provider's reference gives you.","bash","curl -s \\\n  -H \"Authorization: Bearer $API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  \"$API_BASE/scraper/list\"","Read the key from the environment, never from the command line directly — your shell history is a plaintext file and it is backed up.",{"type":125,"title":211,"intro":212,"items":213},"Where the key should live","In rough order of how much trouble each one saves you later.",[214,215,216,217],"In a secret manager your deployment reads at boot, if you have one","In an environment variable set by your orchestrator, if you do not","In a gitignored .env file for local development only, with a .env.example that documents the names but holds no values","Never in the repository, including in a test fixture, including in a commented-out line, including in a notebook you were going to delete",{"type":120,"variant":121,"title":219,"text":220},"Scope and rotate before you need to","If the provider lets you mint more than one key, mint one per environment and one per integration. The entire value of that is on the day you have to revoke one: a single shared key means revoking it takes down everything at once, and you will discover which systems used it by finding out what broke. Rotating is also the only cheap way to recover from a key that has leaked into a log you did not know existed.",{"type":79,"title":222,"paragraphs":223},"Reading the first response properly",[224,225,226],"Two habits here pay for themselves within a week.","First, before you write any parsing code, print the whole response once and read it. Not the field you came for — all of it. Data APIs routinely return metadata alongside the payload: a run identifier, a count, a timestamp, a flag saying the result was truncated. The truncation flag in particular is the kind of thing you find out about by missing it.","Second, check the status code and the body separately. A great many APIs return 200 with an error object inside, because the HTTP request succeeded even though the thing you asked for did not. If your client only branches on the status code, those failures become empty rows rather than exceptions.",{"type":84,"title":228,"intro":229,"headers":230,"rows":234},"Status codes worth branching on","Different responses, genuinely different handling. Treating them all as \"error\" is why retries get expensive.",[231,232,233],"Code","What it usually means","What to do",[235,239,243,247,251,255],[236,237,238],"200 with an error body","The request was valid, the operation was not","Read the body. Do not retry — it will fail identically.",[240,241,242],"400","Your payload is malformed or a field is unknown","Fix the request. Retrying is pointless.",[244,245,246],"401 / 403","Key missing, wrong, revoked, or out of scope","Stop and alert. A retry loop on a dead key is just noise.",[248,249,250],"404","Wrong path, or an object that no longer exists","Check the reference before assuming the route is gone.",[252,253,254],"429","You are over a rate limit","Back off, honour Retry-After if present, and reduce concurrency.",[256,257,258],"5xx","Their problem, possibly transient","Retry with exponential backoff and a cap. See lesson five.",{"type":84,"title":260,"intro":261,"headers":262,"rows":266},"Worked example: three responses that are all 200","Branching on the status code alone is the specific mistake this lesson exists to prevent. All three of these came back 200 OK. One of them is a success.",[263,264,265],"Body","What it actually means","What a status-only check does with it",[267,271,275],[268,269,270],"{ \"run_id\": \"r_8812\", \"rows\": [ … 412 rows … ] }","A success","The right thing, by luck",[272,273,274],"{ \"run_id\": \"r_8813\", \"rows\": [] }","The run completed and found nothing — a site change, a dead URL list, or a genuinely empty result","Writes an empty table over yesterday's good one",[276,277,278],"{ \"error\": \"quota_exceeded\", \"retry_after\": 3600 }","You are out of credit","Records a success, stores no rows, and tells nobody",{"type":125,"title":165,"intro":280,"items":281},"Authentication and the first call are where habits get set that are painful to change a year later.",[282,283,284,285,286,287],"Branching on the status and not the body. Two of the three rows above are failures wearing a 200.","Putting the key in the query string. It lands in server logs, in browser history, in referrer headers and in every screenshot of the address bar — none of which rotate when you do.","One key for everything. When it leaks, and eventually one does, revoking it takes down production, the staging job and somebody's notebook at the same moment.","Not establishing the billable unit on day one. Requests, pages fetched and rows returned are three different numbers, and a quote will use whichever is smallest.","Testing with a key that has more scope than production will have. The call works in development and 403s on deploy, which is the worst possible time to find out.","Retrying a 401. A second attempt changes nothing about a wrong credential, and some providers count the attempts against you.",{"type":79,"title":289,"paragraphs":290},"One thing to verify on day one",[291,292],"Find out, explicitly, what counts as a billable unit. Is it a request you make, a page the provider fetches, or a row you receive? These are not the same number, and the gap between them is where surprise invoices come from.","A run that fetches two thousand pages and returns twelve hundred rows has been billed for two thousand on most pricing models, because the fetch is the cost. If you are budgeting on rows you will be wrong by whatever your failure rate is — and your failure rate is a property of the sites you chose, not of the vendor.",[294,295,296,297],"Use the Authorization header, not a query-string key. Query strings leak into logs, history and referrers.","One key per environment and per integration, so that revoking one does not take everything down.","A 200 can contain an error. Branch on the body as well as the status.","Establish what a billable unit is on day one: requests, fetched pages and returned rows are three different numbers.",{"text":299,"label":300,"url":301},"Next: the step most people skip, and the reason their results come back missing half the fields.","Lesson 3: declaring the fields you want","/learn/product-data-api/declare-the-fields-you-want",{"slug":303,"nav_title":304,"title":305,"summary":306,"time":307,"needs_account":308,"seo":309,"blocks":313,"takeaways":397,"next_step":402},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.","12 min",true,{"title":310,"description":311,"keywords":312},"Declaring an Extraction Schema: Why Your API Fields Are Empty","How to declare the fields a product data API should return, why types and required lists matter, and the common mistake that makes a column come back null on every row.","extraction schema, json schema api, product data fields, api returns null, scraper schema design, structured data extraction",[314,321,325,331,337,340,349,355,383,392],{"type":120,"variant":315,"title":316,"text":317,"cta":318},"product","This lesson uses Scrapewise","The principle is general — every extraction API has some version of a declared schema — but the specific behaviour described here, particularly the required-list rule, is ours. If you are using something else, read it as a prompt to go and test the same question against your own provider, because the answer is rarely in the documentation.",{"label":319,"url":320},"Start with 5 free requests","https://portal.scrapewise.ai/register",{"type":79,"paragraphs":322},[323,324],"A modern extraction API does not have a fixed response shape. You tell it what a product looks like to you, and it goes and finds those things on the page. That is enormously more flexible than a fixed endpoint, and it moves a responsibility onto you that fixed endpoints never did: if you describe the fields badly, you get bad fields, and nothing anywhere will tell you that is what happened.","The schema is a JSON Schema document. Each field has a name, a type, and usually a description. The names are yours — call the price field price or unit_cost or whatever your downstream job already expects, because renaming it later means touching every consumer.",{"type":204,"title":326,"intro":327,"language":328,"code":329,"caption":330},"A minimal product schema","Six fields, which is more than most people need and fewer than most people declare.","json","{\n  \"type\": \"object\",\n  \"properties\": {\n    \"product_title\": { \"type\": \"string\" },\n    \"price\":         { \"type\": [\"number\", \"null\"] },\n    \"currency\":      { \"type\": [\"string\", \"null\"] },\n    \"availability\":  { \"type\": [\"string\", \"null\"] },\n    \"brand\":         { \"type\": [\"string\", \"null\"] },\n    \"sku\":           { \"type\": [\"string\", \"null\"] }\n  },\n  \"required\": [\n    \"product_title\", \"price\", \"currency\",\n    \"availability\", \"brand\", \"sku\"\n  ]\n}","Note that every property is also in required, including the ones that are allowed to be null. That is not a contradiction — see below.",{"type":79,"title":332,"paragraphs":333},"The required list is what makes a field appear",[334,335,336],"This is the single most expensive thing to learn by accident, so here it is plainly: listing a field in required is what makes the extractor emit it. A property that is declared but not required can be omitted entirely, and when it is omitted you do not get a null — you get no key at all, on every row, forever.","The intuition most developers bring is the validation intuition: required means this must not be missing, so leave optional fields out and let them be absent when the page does not have them. That reasoning produces a feed with four columns when you asked for nine, and the symptom looks exactly like \"the site does not publish brand\" rather than \"we never asked for brand\".","The fix is to put every field in required and allow null in its type. You are then saying two different things in the two places: the type says this value may legitimately be absent, and the required list says the key must be present regardless. That is the combination that gives you a stable column set, which is what a warehouse table needs.",{"type":120,"variant":121,"title":338,"text":339},"Field descriptions are not instructions","It is tempting to write a long description explaining what you want — \"the regular price before any discount, excluding delivery\". Test whether that actually changes anything before you rely on it. On our extractor it does not; the required list drives emission and the description is documentation for humans. Writing careful prose into a field nobody reads is a satisfying way to spend an afternoon and change nothing.",{"type":125,"title":341,"intro":342,"items":343},"Type choices that save a conversion later","Small decisions, each of which costs an hour downstream if you get it wrong.",[344,345,346,347,348],"Numbers should be numbers, with no currency symbol and no thousands separator. A price that arrives as the string \"1.299,00 €\" is now your parsing problem, and the comma means different things in different markets.","Use [\"number\", \"null\"] rather than [\"string\", \"null\"] for anything you will do arithmetic on, even if the page shows it with a unit.","Keep currency as its own field rather than inferring it from the domain. Multi-market storefronts exist and they will catch you out.","Availability is better as the raw string the page published than as a boolean you guessed. \"Ships in 2–3 weeks\" is not in stock and it is not out of stock.","Always carry the source URL on the row. Every single data question you will be asked begins with \"where did this come from\".",{"type":79,"title":350,"paragraphs":351},"Declare fewer fields than you think you need",[352,353,354],"There is a real cost to a wide schema. More fields means more for the extractor to look for, which means more latency and, on some pricing models, more cost per page. It also means more columns that can be empty, and an empty column is not neutral — it is something a colleague will eventually interpret as a zero.","A better pattern is to scrape only the raw values and compute anything derived afterwards. Unit price, price in a common currency, discount percentage: none of those should be fields the extractor hunts for. They are arithmetic on fields you already have, and doing them after collection means a currency rate change does not require a re-run.","There is a mechanical reason for this too. If a derived value is declared as a scraped field, a post-processing rule that wants to write to that same name usually cannot — the scraped column occupies it. So the derived column stays empty and the rule looks broken when actually it was blocked.",{"type":84,"title":356,"intro":357,"headers":358,"rows":362},"Worked example: the same schema, two required arrays","This is the single behaviour that catches most people, and it is worth seeing side by side. Both schemas declare the same five properties. The only difference is what is listed as required — and that difference decides whether a column exists at all. The left-hand column is how a spreadsheet ends up with a header row that moves between runs; the right-hand one is a stable column set with honest empties.",[359,360,361],"Field","required is [\"name\", \"price\"]","required is all five, types allow null",[363,366,369,373,375,379],[364,365,365],"name","\"Acme Widget 500g\"",[367,368,368],"price","19.99",[370,371,372],"sale_price, no promotion on the page","key absent from the row entirely","null",[374,371,372],"pack_size, not printed on the page",[376,377,378],"currency, present on the page","sometimes there, sometimes not","\"EUR\"",[380,381,382],"Columns in the resulting table","varies row by row","five, every row, every run",{"type":125,"title":165,"intro":384,"items":385},"Schema mistakes do not raise errors. They produce a table that is a little bit wrong in a way that is hard to see.",[386,387,388,389,390,391],"Declaring a field and leaving it out of the required array. It is described, it is understood, and it is not emitted — the quietest failure in this whole course.","Writing instructions into the field description. A description says what the field is, not what to do with it; arithmetic, conversions and clean-up all belong downstream.","Declaring a field whose name collides with something computed later. A scraped unit_price column blocks the rule meant to produce one, and the rule then simply never runs.","Asking for thirty fields on the first pass. Each one is another thing that can be wrong, and twenty-five of them will never be read by anybody.","Typing a number as a string because the page shows \"500 g\". Keep the raw text in its own field and let the number field be a number or null, otherwise you are parsing on every read forever.","Scaling to ten thousand pages without reading one complete row. A missing column is obvious in a single row and invisible in a summary count.",{"type":79,"title":393,"paragraphs":394},"Verify with one page before you run ten thousand",[395,396],"Point the configuration at a single known URL and look at the row. Not the row count — the row. Check that every key you declared is present, that the price is a number, that the currency is what you expected, and that the title is the product rather than the category.","If a column is missing, the schema is the first place to look, not the site. If a column is present but null on a page where you can see the value with your own eyes, that is a genuine extraction problem and worth reporting. Those two failures look identical in a dashboard and have completely different fixes.",[398,399,400,401],"Listing a field in the required array is what makes the extractor emit it. Declared-but-not-required fields come back as no key at all.","Put every field in required and allow null in its type. That gives a stable column set with honest empties.","Scrape raw values only. Compute unit prices, conversions and discounts afterwards, or the derived column gets blocked by the scraped one.","Verify against one page and read the whole row before scaling to thousands.",{"text":403,"label":404,"url":405},"Next: why the call that starts a run does not return your data, and how to work with that properly.","Lesson 4: asynchronous runs and polling","/learn/product-data-api/asynchronous-runs-and-polling",{"slug":407,"nav_title":408,"title":409,"summary":410,"time":411,"needs_account":72,"seo":412,"blocks":416,"takeaways":529,"next_step":534},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.","11 min",{"title":413,"description":414,"keywords":415},"Asynchronous API Runs: Polling, Job State and Partial Results","Why web data APIs return a job id instead of data, how to poll correctly with backoff, how to handle pagination, and what to do when a run completes partially.","asynchronous api, polling api job status, api pagination, long running request, webhook vs polling, partial results",[417,422,436,442,445,451,457,481,520],{"type":79,"paragraphs":418},[419,420,421],"The first thing that surprises people coming from ordinary REST is that starting a collection does not give you any data. It gives you an identifier. The work happens somewhere else, over minutes or hours, and you come back for the result.","This is not an architectural preference. Fetching a single page takes seconds in the best case, and much longer when the site is slow, when a retry is needed, or when a headless browser has to render it. A synchronous endpoint over ten thousand URLs would hold a connection open for hours, and the first network blip would lose the entire result. So the API hands you a job.","The consequence for your code is that there are now three distinct operations where you expected one: start, check, collect.",{"type":423,"title":424,"intro":425,"items":426},"steps","The three-step shape","Every asynchronous data API is some variation of this, whatever the endpoints are called.",[427,430,433],{"title":428,"text":429},"Start the run","Post the configuration, or the identifier of a saved configuration, and get back a job or run id. Store that id somewhere durable immediately. If your process dies here and the id was only in memory, you have paid for a run you cannot collect.",{"title":431,"text":432},"Poll until it settles","Ask for the run's state on an interval until it reaches a terminal state. Terminal usually means finished, failed or stopped. Anything else means keep waiting.",{"title":434,"text":435},"Collect the rows","Fetch the results, usually paginated. This is a separate call and often a separate rate limit, which matters because collecting a large run can take longer than you expect.",{"type":79,"title":437,"paragraphs":438},"How to poll without being rude",[439,440,441],"A fixed one-second poll on a job that takes forty minutes is two and a half thousand pointless requests, and on some providers it is two and a half thousand requests against your rate limit.","Use exponential backoff with a ceiling. Start at a few seconds, double each time, cap at thirty or sixty seconds. For a job you expect to run for an hour you will make a few dozen calls rather than thousands, and you will still notice completion within a minute of it happening.","If the provider offers a webhook, use it and keep the polling as a fallback rather than deleting it. Webhooks get lost — a deploy restarts your process mid-delivery, a firewall rule changes, the retry policy is less generous than you assumed. A slow poll behind a fast webhook costs almost nothing and converts a silent stall into a late notification.",{"type":120,"variant":121,"title":443,"text":444},"Set a deadline, and decide now what happens when it passes","A run that never reaches a terminal state is the failure mode that hurts most, because nothing errors. Your poller just keeps politely asking. Pick a maximum wall-clock time based on the size of the run, and when it elapses, stop the run explicitly rather than abandoning it — an abandoned run on most platforms carries on working and carries on billing. The decision to make in advance is whether a timed-out run's partial rows are usable or discarded, because making that call at three in the morning goes badly.",{"type":79,"title":446,"paragraphs":447},"Pagination and the row cap nobody mentions",[448,449,450],"Results come back in pages. The mechanics are ordinary — a limit, an offset or a cursor, a loop — but there are two traps.","The first is that sample and preview endpoints frequently cap out at a round number like a hundred rows. If you test with one of those and then build your pagination against it, you will conclude the loop works and that runs are smaller than they are. Check whether the endpoint you are looping over is the full-results endpoint or the preview one.","The second is that the count you use to decide when to stop should come from the run's own metadata, not from counting what you have received. If a page comes back short because of a transient error and your loop uses received-count as its terminator, you will stop early and treat a partial collection as a complete one.",{"type":79,"title":452,"paragraphs":453},"Partial results are the normal case",[454,455,456],"A run over several thousand URLs will not return several thousand rows. Some pages will be gone, some will be a redirect to a category listing, some will be behind a bot wall on the day you ran it. A ninety per cent return is a good run.","So completion and completeness are different questions, and your code should ask both. The run finished — fine. Did it return roughly what the last one did? That second question is the one that catches real problems, and lesson five of course one covers how to baseline it.","The thing not to do is treat a short run as a failure and retry the whole thing. You will pay twice to collect the same ninety per cent, and the ten per cent that failed will mostly fail again, because the reason was the page and not the attempt.",{"type":84,"title":458,"intro":459,"headers":460,"rows":464},"Terminal states and what each one means for your data","Names vary. The categories do not.",[461,462,463],"State","Meaning","Are the rows usable?",[465,469,473,477],[466,467,468],"Finished","The run processed its whole input list","Yes, subject to the usual per-page failures",[470,471,472],"Stopped","Something ended it early — you, a cap, or a deadline","Usually yes, but the input was not fully covered",[474,475,476],"Failed","The run itself could not proceed, often a config error","Treat as none; fix the configuration first",[478,479,480],"Finished with errors","Completed, but a meaningful share of pages did not return","Yes, and the error breakdown is the thing to read",{"type":84,"title":482,"intro":483,"headers":484,"rows":489},"Worked example: a polling schedule that does not hammer","Exponential backoff with a thirty-second cap, against a run that takes four minutes. Twelve requests instead of the 270 a one-second poll would have made, and the longest you ever sit on a finished run is thirty seconds.",[485,486,487,488],"Poll","Wait before it","Elapsed","Status",[490,495,499,503,507,511,515],[491,492,493,494],"1","2s","0:02","running",[496,497,498,494],"2","4s","0:06",[500,501,502,494],"3","8s","0:14",[504,505,506,494],"4","16s","0:30",[508,509,510,494],"5","30s (cap reached)","1:00",[512,513,514,494],"6 to 11","30s each","4:00",[516,517,518,519],"12","30s","4:30","succeeded",{"type":125,"title":165,"intro":521,"items":522},"Asynchronous collection fails in ways that look like success, which is why the counts matter more than the state.",[523,524,525,526,527,528],"Polling on a fixed one-second interval. That is 270 requests where twelve would do, and on a provider with a global rate limit the poller can starve the collection it is waiting for.","Not storing the run id before the first poll. The process restarts, the id is gone, and the run carries on — and carries on billing — with nobody collecting it.","No wall-clock deadline. A stalled run polls forever, and forever is a line on an invoice.","Treating finished as complete. A run can terminate successfully having fetched 900 of 1,200 pages. Ask for the state and the counts, and compare the counts against the last good run.","Ignoring the row cap. Many endpoints return at most a few hundred rows per page, so a run that produced 4,000 rows hands back 100 and a cursor, and a naive client records 100.","Throwing away partial results on failure. Eighty per cent of a run is usually worth keeping, provided the gap is recorded rather than quietly absorbed into an average.",[530,531,532,533],"Start, poll, collect. Store the run id durably the moment you get it.","Poll with exponential backoff and a cap. Use webhooks if offered, but keep a slow poll as a fallback.","Set a wall-clock deadline and stop a stalled run explicitly — abandoned runs keep billing.","Partial results are normal. Completion and completeness are two different questions and you should ask both.",{"text":535,"label":536,"url":537},"Next: retrying properly, including how to avoid paying twice for the same page.","Lesson 5: errors, retries and double billing","/learn/product-data-api/errors-retries-and-double-billing",{"slug":539,"nav_title":540,"title":541,"summary":542,"time":411,"needs_account":72,"seo":543,"blocks":547,"takeaways":650,"next_step":655},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"title":544,"description":545,"keywords":546},"API Retries Without Double Billing: Idempotency and Backoff","Which API errors to retry and which to stop on, how idempotency keys prevent duplicate charges, exponential backoff with jitter, and rate limit handling.","api retry strategy, idempotency key, exponential backoff, rate limit 429, duplicate api charge, api error handling",[548,552,582,587,593,599,602,611,635,644],{"type":79,"paragraphs":549},[550,551],"Retry logic is where a tidy integration becomes an expensive one. The default instinct — wrap the call in a loop, try three times, move on — is wrong in two directions at once. It retries things that will never succeed, and it retries things that already succeeded.","The second of those is the one that costs money. A request can time out on your side after the server has already done the work. From your perspective nothing happened; from the provider's, a run started and a bill was incurred. Retry and you have two.",{"type":84,"title":553,"intro":554,"headers":555,"rows":559},"Retry or stop","The useful split is not by status code but by whether trying again could plausibly change the outcome.",[556,557,558],"Failure","Retry?","Why",[560,564,568,571,575,578,580],[561,562,563],"Connection reset, DNS blip, timeout","Yes, with backoff","Genuinely transient. The next attempt is a different network moment.",[565,566,567],"500, 502, 503, 504","Yes, with backoff and a cap","Their side, usually brief. Cap it so a long outage does not become an infinite loop.",[252,569,570],"Yes, but slower","Honour Retry-After. Also reduce concurrency — retrying at the same rate just re-earns the 429.",[572,573,574],"400, 422","No","The payload is wrong. It will be wrong on the next attempt too.",[576,573,577],"401, 403","The key is missing, revoked or unscoped. Alert a human instead.",[248,573,579],"Check the path against the reference. Retrying a typo is just a slower typo.",[236,573,581],"Deterministic. Read the body and handle it as the failure it is.",{"type":79,"title":583,"paragraphs":584},"Backoff needs jitter",[585,586],"Exponential backoff alone has a well-known failure mode. If fifty of your workers all fail at the same moment — which is exactly what happens during a brief provider outage — they all back off by the same amount and all retry at the same instant. You have rebuilt the thundering herd with extra steps, and your retry storm is what keeps the recovery from happening.","Add randomness. Wait somewhere between zero and the backoff interval rather than exactly the interval. It is a one-line change and it is the difference between a recovery and a second outage.",{"type":204,"title":588,"intro":589,"language":590,"code":591,"caption":592},"Backoff with jitter, in the smallest honest form","Deliberately boring. The important parts are the cap, the jitter, and that it only loops on retryable failures.","python","import random, time\n\ndef with_retries(call, attempts=5, base=2.0, cap=60.0):\n    for n in range(attempts):\n        try:\n            return call()\n        except Retryable:\n            if n == attempts - 1:\n                raise\n            delay = min(cap, base * (2 ** n))\n            time.sleep(random.uniform(0, delay))","Retryable is your own exception, raised only for the rows in the table above that say yes. Everything else should propagate immediately.",{"type":79,"title":594,"paragraphs":595},"Idempotency keys",[596,597,598],"An idempotency key is a value you generate and attach to a request that starts work. If the provider supports it, a second request carrying the same key does not start a second run — it returns the result of the first one.","This is the clean answer to the timed-out-but-succeeded problem. Generate the key before the first attempt, reuse it on every retry of that same logical operation, and a retry becomes safe rather than expensive. Generating a fresh key per attempt defeats the entire mechanism, which is a surprisingly common mistake because the generation often lives inside the retry loop by accident.","Two practical notes. Support is usually opt-in, so it does nothing unless you send the header. And it is often supported on more methods than you would guess, including deletes, which matters because a timed-out delete is just as ambiguous as a timed-out create.",{"type":120,"variant":121,"title":600,"text":601},"Rate limits are frequently global, not per-endpoint","It is natural to assume that a heavy endpoint has its own budget and that a cheap one, like checking a job's status, does not count. Often untrue. If the limit is applied globally, an aggressive poller can consume the allowance your actual collection needs, and the symptom is that the collection starts failing while the polling keeps working perfectly. If you are seeing 429s, count all your calls, not just the obvious ones.",{"type":125,"title":603,"intro":604,"items":605},"Log enough to diagnose it later","What you want in the log line at the moment of failure, because none of it is recoverable afterwards.",[606,607,608,609,610],"The run or request id the provider gave you — without it, support cannot help and neither can you","The status code and the first part of the response body, not just the exception message","Which attempt number this was, so a one-off is distinguishable from a systematic failure","The timestamp in UTC, because the provider's logs are in UTC and correlating across zones at 3am is how mistakes happen","A durable note of whether work may have started, so the next operator knows whether a retry is safe",{"type":84,"title":612,"intro":613,"headers":614,"rows":619},"Worked example: how one timeout becomes three invoices","A start-run request times out at the client after thirty seconds. The server received it and is working on it. Without an idempotency key every retry is a new run, and all of them finish. With one key generated before the first attempt, attempts two and three return run A's id and the bill is one run — but the key has to exist before attempt one. A key generated inside the retry loop is a new key each time and does nothing whatsoever.",[615,616,617,618],"Attempt","What the client sees","What the server does","Runs in flight",[620,623,626,630],[491,621,622,491],"timeout after 30s","accepted, starts run A",[624,621,625,496],"2, a retry","accepted, starts run B",[627,628,629,500],"3, a retry","200, run id C","accepted, starts run C",[631,632,633,634],"Outcome","one run id","three complete runs","three times the pages, a third of them collected",{"type":125,"title":165,"intro":636,"items":637},"Retry logic is written on a good day and exercised on a bad one, which is why these are so common.",[638,639,640,641,642,643],"Retrying a 400. Nothing about a second attempt changes a malformed request, and some providers charge for the attempt anyway.","Backoff without jitter. Every client that failed during the outage retries on exactly the same schedule and rebuilds the stampede at the moment the service is trying to recover.","Generating the idempotency key inside the retry loop. A fresh key per attempt is identical to no key, and it reads as correct in the code.","Assuming the rate limit is per endpoint. It is frequently global, so an aggressive poller on one route can 429 the route that actually matters.","Retrying a quota or payment error. That is the class that means stop: the system is working correctly and telling you that you are out of credit.","Logging that a call failed without logging what was sent. A month later you have a count of failures and no way to reproduce a single one.",{"type":79,"title":645,"paragraphs":646},"The failure that means stop, not try harder",[647,648,649],"There is one category where retrying is not just useless but actively the wrong instinct: when the page is refusing you rather than failing.","A bot wall that has decided your traffic is automated does not give you a flaky response that works on the third attempt. It gives you a consistent refusal, and hammering it makes the classification worse. Three retries against a wall is three times the cost for the same answer, plus a stronger signal that you are what the wall thinks you are.","This is the main subject of the next course, and the distinction worth carrying into it is between a page that failed and a page that declined. Your retry logic can only help with the first.",[651,652,653,654],"Split failures by whether a second attempt could plausibly change the outcome, not by status code family.","Exponential backoff without jitter rebuilds the thundering herd during exactly the outage you are trying to survive.","Generate the idempotency key before the first attempt and reuse it on every retry, or it does nothing.","Rate limits are often global. An aggressive poller can starve the collection it is polling for.",{"text":656,"label":657,"url":658},"Next: getting the feed out of the API and into something your team actually queries.","Lesson 6: putting the feed into your stack","/learn/product-data-api/put-the-feed-into-your-stack",{"slug":660,"nav_title":661,"title":662,"summary":663,"time":307,"needs_account":308,"seo":664,"blocks":668,"takeaways":772,"next_step":777},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",{"title":665,"description":666,"keywords":667},"Loading a Product Price Feed Into a Warehouse Without Drift","How to schedule collection, design the landing table, keep price history, and detect the day a feed silently halves. Practical patterns for a durable data pipeline.","price feed warehouse, product data pipeline, data loading schema, price history table, feed monitoring, etl product data",[669,672,676,682,702,708,717,720,726,758,767],{"type":120,"variant":315,"title":316,"text":670,"cta":671},"The scheduling and export mechanics described are ours. The table design and the monitoring are not vendor-specific and are the part worth copying regardless of where your data comes from.",{"label":319,"url":320},{"type":79,"paragraphs":673},[674,675],"The integration is working, the rows come back, and now the real question: where do they go, and will the answer still be right in six months?","Most feeds do not die dramatically. They drift. A column that used to be populated goes mostly null and nobody notices because nobody queries it. A site is dropped from the schedule during a cleanup and the dashboard quietly covers one fewer competitor. The table has a last_updated column that stopped updating. All three are invisible if the only thing you check is whether the pipeline ran.",{"type":79,"title":677,"paragraphs":678},"Decide the cadence from the decision, not the data",[679,680,681],"The common instinct is to collect as often as the provider allows. That is backwards. The right frequency is set by how often somebody acts on the number.","If your repricing job runs once a day at six in the morning, collecting hourly buys you nothing except twenty-three discarded datasets and a bill. If a buyer checks a dashboard on Monday mornings, daily collection is already generous. Hourly makes sense for a narrow set of volatile, high-value SKUs, and almost never for the whole catalogue.","Set the collection to finish comfortably before the thing that consumes it starts. Comfortably means with enough margin that a slow run does not mean a stale decision — and a slow run is normal, because run duration is a property of how the sites behaved that day.",{"type":84,"title":683,"intro":684,"headers":685,"rows":689},"Three ways to get the rows out, and when each fits","Most platforms offer all three. They are not interchangeable.",[686,687,688],"Method","Good for","The catch",[690,694,698],[691,692,693],"Scheduled export to a file","Warehouse loads, BI tools, anything batch","You need a loader on the other end, and somebody has to notice when a file does not arrive",[695,696,697],"Pull the results over the API","Full control, custom transforms, backfills","Pagination and retry logic are now yours, and so is the scheduling",[699,700,701],"Webhook on completion","Triggering a downstream job the moment data is ready","Deliveries get lost. Pair with a poll, as in lesson four.",{"type":79,"title":703,"paragraphs":704},"Land it raw, then transform",[705,706,707],"Write what the API returned into a landing table, unmodified, with a collection timestamp and the run id. Then build your cleaned view on top of it. This is standard advice and it is standard for a reason: the day somebody asks why Tuesday's price looks wrong, you want to be able to see what actually arrived on Tuesday rather than what your transform made of it.","It also makes a transform bug recoverable. If the cleaned table is the only copy and your parser mangled a currency for a fortnight, that fortnight is gone. If the landing table is intact you re-run the transform and the fortnight comes back.","Keep the source URL on every row through every layer. It is the only thing that lets a human verify a number by opening a page, and it will be asked for.",{"type":125,"title":709,"intro":710,"items":711},"Columns worth having that people leave out","Each of these exists to answer a question somebody will eventually ask.",[712,713,714,715,716],"collected_at — when the value was read, not when the row was written. These differ and the difference is the age of your data.","run_id — so a suspicious batch can be traced to one run and compared against its neighbours","source_url — the exact page, so a number can be verified by a human in one click","raw_price_text — what the page actually said, before parsing. Invaluable the first time a decimal separator goes wrong.","A row per observation rather than an updated row per product, so that price history exists by construction rather than being added later at great expense",{"type":120,"variant":121,"title":718,"text":719},"Overwriting is how you lose history you did not know you needed","An update-in-place table answers \"what is the price now\" and nothing else. The first interesting question anyone asks of a price feed is \"when did they change it\", and if you have been overwriting, the honest answer is that you do not know and cannot find out. Append-only costs more storage than you will ever notice and it is the single decision that most determines whether this dataset is useful in a year.",{"type":79,"title":721,"paragraphs":722},"Monitor the shape of the data, not the health of the job",[723,724,725],"A green pipeline run tells you the code did not throw. It says nothing about whether the data is right, and the failures that matter are almost all of the second kind.","Three checks catch most of it. Row count against the trailing average, per site, with an alert on a meaningful drop — a feed halving is a far more common failure than a feed stopping. Null rate per column against its own baseline, because a column going empty is a redesign signal. And a freshness check that actually compares collected_at to now, rather than trusting that the scheduler ran.","The reason to do this per site rather than in aggregate is that aggregate numbers hide exactly the failures you care about. Forty sites holding steady and one going to zero is a two and a half per cent drop in the total, which no sensible threshold will catch, and it is also one competitor having completely vanished from your pricing decisions.",{"type":84,"title":727,"intro":728,"headers":729,"rows":733},"Worked example: what update-in-place destroys","Four observations of one SKU, stored two ways. The left-hand table is what almost everyone builds first, because it matches how a price feels — a current value. The right-hand one answers the question that always gets asked. The cost of the right-hand version is 365 rows a year per SKU per competitor, which for 1,200 SKUs and five competitors is about 2.2 million rows: unremarkable for any database built this decade, and more than any spreadsheet should be asked to hold.",[730,731,732],"Observation","Update-in-place","Append-only",[734,738,741,745,748,752,755],[735,736,737],"Monday, €229.00","one row: €229.00","row 1",[739,736,740],"Tuesday, €229.00","row 2",[742,743,744],"Wednesday, €189.00","one row: €189.00","row 3",[746,743,747],"Thursday, €189.00","row 4",[749,750,751],"\"When did they drop it, and from what?\"","unanswerable","Wednesday, from €229.00",[753,750,754],"\"How long has the promotion run?\"","two days so far",[756,491,757],"Rows after a year of daily runs, one SKU","365",{"type":125,"title":165,"intro":759,"items":760},"Six months in, a feed is either still trusted or quietly worked around. These are the decisions that determine which.",[761,762,763,764,765,766],"Updating in place. The cheapest decision on day one and the most regretted, because the history cannot be reconstructed afterwards.","Transforming before landing. If the raw response is discarded, a bug in the transform costs you the data instead of an afternoon of reprocessing.","Setting cadence from what the provider allows. Set it from how often somebody actually changes a price; everything above that is pages you pay for and nobody reads.","Monitoring the job instead of the shape. Row count and null rate per site against a trailing baseline catch the degradations. A green job catches almost nothing.","Monitoring in aggregate. One competitor out of five dropping to zero moves the total by a fifth, which sits comfortably under any threshold you would have set.","Leaving the coverage gaps undocumented. The feed does not include third-party marketplace sellers, or a market nobody configured, or anything behind a login — and the first person surprised by that will be the one presenting from it.",{"type":79,"title":768,"paragraphs":769},"Write down what the feed does not cover",[770,771],"Every feed has holes: sites that block, pages that do not publish a price, categories that were never in scope. These are fine. What is not fine is that they live in one engineer's head.","Keep a short document next to the pipeline that lists which sites are in scope, which are excluded and why, and which fields are known to be unreliable where. It takes twenty minutes to write and it prevents the single most damaging thing a data feed can do, which is to be quietly interpreted as complete by somebody who was not there when it was built.",[773,774,775,776],"Set collection cadence from how often someone acts on the number, not from what the provider allows.","Land the raw response first, transform on top. A transform bug is then recoverable; otherwise the data is gone.","Append a row per observation. Update-in-place destroys the price history that is the first thing anyone asks for.","Monitor row count and null rate per site against a baseline. Aggregate monitoring hides the one competitor that vanished.",{"text":778,"label":779,"url":780},"That is the course. The next one is about the thing that breaks this feed most often: the sites themselves changing, and occasionally deciding they would rather not be read.","Course four: keeping scrapers alive","/learn/keep-scrapers-alive",[782,834,874,882,921,959],{"order":783,"slug":784,"title":785,"subtitle":786,"cardText":787,"level":788,"time":789,"lessonCount":790,"lessons":791},1,"competitor-price-monitoring","Build a competitor price monitoring pipeline","Price monitoring looks like a scraping problem for about a week. Then you discover that scraping was the easy part, and the project actually lives or dies on which competitors you picked, whether their listings are really the same product as yours, and whether anyone notices the morning the feed comes back half empty. This course is those eight decisions, in the order you have to make them.","From \"we check three competitors by hand on Mondays\" to a feed you trust enough to reprice from. The eight decisions in order, including the two that quietly ruin most projects.","No coding required","8 lessons, about 90 minutes",8,[792,798,803,808,813,819,824,829],{"slug":793,"navTitle":794,"title":795,"summary":796,"time":797,"needsAccount":72},"what-is-competitor-price-monitoring","What it actually is","What competitor price monitoring actually is","The four stages of a price pipeline, why only two of them are scraping, and the one question to ask before you build anything.","9 min",{"slug":799,"navTitle":800,"title":801,"summary":802,"time":411,"needsAccount":72},"choose-competitors-and-skus","Choosing what to track","Choosing which competitors and which SKUs to track","How to build a list that is small enough to afford and large enough to matter, using margin at risk rather than gut feel.",{"slug":804,"navTitle":805,"title":806,"summary":807,"time":307,"needsAccount":308},"find-competitor-product-urls","Finding product URLs","Finding every competitor product URL without copying them by hand","Four ways to get a competitor's full product URL list, ranked by how much work they are, and what to do when none of them work.",{"slug":809,"navTitle":810,"title":811,"summary":812,"time":307,"needsAccount":308},"extract-price-stock-and-shipping","Extracting the fields","Getting price, stock and shipping off the page","Which fields to extract, why the sale price is two fields and not one, and the four ways a price appears on a page.",{"slug":814,"navTitle":815,"title":816,"summary":817,"time":818,"needsAccount":72},"match-listings-to-your-catalogue","Matching to your catalogue","Matching competitor listings to your own catalogue","The stage that decides whether your feed is intelligence or fiction, and the denominator trick that makes bad match rates look good.","13 min",{"slug":820,"navTitle":821,"title":822,"summary":823,"time":411,"needsAccount":72},"schedule-runs-and-catch-silent-failure","Scheduling and data quality","Scheduling runs and catching silent data loss","How often to actually check, and the four alerts that catch a degrading feed before someone reprices from it.",{"slug":825,"navTitle":826,"title":827,"summary":828,"time":797,"needsAccount":308},"export-to-sheets-bi-and-erp","Getting the data out","Getting the data into Sheets, BI or your ERP","Four delivery routes ranked by how likely they are to actually get used, and the column contract that stops downstream jobs breaking.",{"slug":830,"navTitle":831,"title":832,"summary":833,"time":307,"needsAccount":72},"turn-price-data-into-repricing-rules","From data to decisions","Turning price data into repricing decisions","Why \"match the cheapest\" destroys margin, what a rule needs besides a competitor price, and how to start without automating anything.",{"order":835,"slug":836,"title":837,"subtitle":838,"cardText":839,"level":840,"time":841,"lessonCount":842,"lessons":843},2,"ai-agent-web-data-mcp","Give your AI agent live web data via MCP","Ask an assistant what a product costs today and you will usually get a number. It is often wrong, and it is always wrong in the same way: the model is reconstructing a plausible price from training data rather than looking at a page. This course is about closing that gap properly — what the Model Context Protocol actually is, how to wire a server into a client, how to design tools a model can use without hand-holding, and what to put in place before an agent spends your money.","Your agent is confidently wrong about prices because it has never seen one. What MCP is, how to connect a server, how to design tools a model can actually use, and the guardrails you need before you let it loose.","Comfortable editing a config file","6 lessons, about 60 minutes",6,[844,849,854,859,864,869],{"slug":845,"navTitle":846,"title":847,"summary":848,"time":797,"needsAccount":72},"what-is-mcp","What MCP is","What MCP actually is, in plain terms","The Model Context Protocol described without jargon: what problem it solves, its three primitives, and when it is the wrong tool.",{"slug":850,"navTitle":851,"title":852,"summary":853,"time":71,"needsAccount":72},"why-agents-get-live-data-wrong","Why agents get it wrong","Why your agent's answer about a price is wrong","Four distinct failure modes that all look identical from the outside, and how to tell which one you have before you try to fix it.",{"slug":855,"navTitle":856,"title":857,"summary":858,"time":71,"needsAccount":72},"connect-an-mcp-server","Connecting a server","Connecting an MCP server and proving it works","The config for local and remote servers, the four things that go wrong, and how to verify the tools registered rather than assuming.",{"slug":860,"navTitle":861,"title":862,"summary":863,"time":411,"needsAccount":72},"design-tools-an-agent-can-use","Designing usable tools","Designing tools an agent can actually use","A connected server is not a useful server. The model only sees your tool names, descriptions and parameter schemas, so those three things are the entire user interface. Here is what makes a tool get called correctly and what makes it get ignored.",{"slug":865,"navTitle":866,"title":867,"summary":868,"time":307,"needsAccount":308},"give-an-agent-a-scraper","Giving an agent a scraper","Giving an agent a real price feed","A worked example. Connect the ScrapeWise MCP server to a client, let the agent read a live scraper's output, and watch where the hand-off between \"the data is right\" and \"the answer is right\" actually breaks.",{"slug":870,"navTitle":871,"title":872,"summary":873,"time":411,"needsAccount":72},"guardrails-cost-and-untrusted-content","Guardrails and cost","Guardrails, cost control and untrusted content","Live web access turns an agent into something that can spend money and read text written by strangers. Neither is a reason not to do it. Both are reasons to put limits in before you need them.",{"order":5,"slug":4,"title":17,"subtitle":18,"cardText":8,"level":6,"time":7,"lessonCount":842,"lessons":875},[876,877,878,879,880,881],{"slug":67,"navTitle":68,"title":69,"summary":70,"time":71,"needsAccount":72},{"slug":189,"navTitle":190,"title":191,"summary":192,"time":71,"needsAccount":72},{"slug":303,"navTitle":304,"title":305,"summary":306,"time":307,"needsAccount":308},{"slug":407,"navTitle":408,"title":409,"summary":410,"time":411,"needsAccount":72},{"slug":539,"navTitle":540,"title":541,"summary":542,"time":411,"needsAccount":72},{"slug":660,"navTitle":661,"title":662,"summary":663,"time":307,"needsAccount":308},{"order":883,"slug":884,"title":885,"subtitle":886,"cardText":887,"level":888,"time":889,"lessonCount":842,"lessons":890},4,"keep-scrapers-alive","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.","You already have something running","6 lessons, about 65 minutes",[891,896,901,906,911,916],{"slug":892,"navTitle":893,"title":894,"summary":895,"time":71,"needsAccount":72},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.",{"slug":897,"navTitle":898,"title":899,"summary":900,"time":307,"needsAccount":72},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.",{"slug":902,"navTitle":903,"title":904,"summary":905,"time":411,"needsAccount":72},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.",{"slug":907,"navTitle":908,"title":909,"summary":910,"time":411,"needsAccount":72},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":912,"navTitle":913,"title":914,"summary":915,"time":411,"needsAccount":72},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":917,"navTitle":918,"title":919,"summary":920,"time":71,"needsAccount":72},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"order":922,"slug":923,"title":924,"subtitle":925,"cardText":926,"level":927,"time":7,"lessonCount":842,"lessons":928},5,"matching-products-across-sites","Match the same product across different sites","A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.","Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.","You have data from more than one site",[929,934,939,944,949,954],{"slug":930,"navTitle":931,"title":932,"summary":933,"time":71,"needsAccount":72},"why-matching-is-the-hard-part","Why matching is hard","Why matching is the hard part","The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.",{"slug":935,"navTitle":936,"title":937,"summary":938,"time":411,"needsAccount":72},"identifiers-first-gtin-ean-mpn","Identifiers first","Identifiers first: GTIN, EAN, UPC and MPN","What each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.",{"slug":940,"navTitle":941,"title":942,"summary":943,"time":307,"needsAccount":72},"when-there-is-no-barcode","No barcode","Matching when there is no barcode","Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.",{"slug":945,"navTitle":946,"title":947,"summary":948,"time":307,"needsAccount":72},"variants-bundles-and-multipacks","Variants and packs","Variants, bundles and multipacks","The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.",{"slug":950,"navTitle":951,"title":952,"summary":953,"time":411,"needsAccount":308},"score-confidence-and-build-a-review-queue","Confidence and review","Confidence scores and a review queue worth using","Why one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.",{"slug":955,"navTitle":956,"title":957,"summary":958,"time":411,"needsAccount":72},"measure-your-match-rate-honestly","Measure it honestly","Measure your match rate honestly","The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.",{"order":842,"slug":960,"title":961,"subtitle":962,"cardText":963,"level":964,"time":965,"lessonCount":922,"lessons":966},"web-scraping-legal-and-ethical","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.","No legal background assumed","5 lessons, about 55 minutes",[967,972,977,982,987],{"slug":968,"navTitle":969,"title":970,"summary":971,"time":71,"needsAccount":72},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.",{"slug":973,"navTitle":974,"title":975,"summary":976,"time":411,"needsAccount":72},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.",{"slug":978,"navTitle":979,"title":980,"summary":981,"time":411,"needsAccount":72},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"slug":983,"navTitle":984,"title":985,"summary":986,"time":411,"needsAccount":72},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.",{"slug":988,"navTitle":989,"title":990,"summary":991,"time":307,"needsAccount":72},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.",1791047866696]