[{"data":1,"prerenderedAt":211},["ShallowReactive",2],{"learn-lesson-product-data-api-declare-the-fields-you-want":3},{"course":4,"lesson":67,"index":178,"outline":179,"prev":209,"next":210},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"product-data-api",3,"Comfortable with HTTP and JSON","6 lessons, about 70 minutes","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Product Data API: A Free 6-Lesson Course for Developers","How to pull product, price and stock data over an API. Authentication, declaring fields, asynchronous runs, pagination, retries without double billing, and putting the feed into your stack. Free, ungated.","product data api, price api, ecommerce api, web scraping api, api key authentication, api documentation, scraper api tutorial, rest api product data","A free developer course on pulling product data over an API","Six written lessons: when an API beats a scraper, auth, declaring fields, async runs, error handling, and shipping the feed into your stack.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course three","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.",{"label":21,"url":22},"Start with lesson one","/learn/product-data-api/when-an-api-beats-a-scraper",{"label":24,"url":25},"See what a run returns","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Decide between an official retailer API, a web data API and writing your own scraper, with reasons you can defend in review","Make an authenticated request and read the response without guessing what the fields mean","Declare a schema so the extractor returns the fields you need rather than the ones it felt like returning","Work with asynchronous runs: start, poll, and handle a job that finishes half-done","Retry a failed request without being billed twice for the same page","Land the result in a warehouse table that does not quietly drift out of date",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Developers who have been told to \"get the competitor prices\" and have discovered the retailer has no public API","Data engineers deciding whether to own the collection layer or buy it","Backend teams wiring a third-party feed into an existing pipeline and wanting to know where it will break","Not written for",[44,45],"Readers looking for a no-code setup — course one covers the same ground without a terminal","Anyone wanting a specific vendor's endpoint reference. This is the shape of the problem; the reference lives in that vendor's docs.",{"title":47,"intro":48},"The six lessons","Lessons one and two are short and assume nothing but curl. From three onwards you will get more out of it with a key in your hand.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions developers ask in the first ten minutes.",[54,57,60,63],{"title":55,"description":56},"Does this teach a specific vendor's API?","No, deliberately. Endpoint names change and a course that hardcodes them rots. What does not change is the shape: authenticate, declare what you want, start a run, poll it, handle partial results, retry safely. Learn the shape and any vendor's reference becomes a lookup rather than a tutorial.",{"title":58,"description":59},"Why is it not just a GET that returns the price?","Because fetching a page takes seconds and fetching ten thousand takes hours. Any API that hides that behind a synchronous request is either very slow or quietly returning you a cached number. Lesson four is about why the asynchronous shape exists and how to work with it instead of around it.",{"title":61,"description":62},"Do I need a Scrapewise account?","Lessons three and six include a Scrapewise-specific walkthrough and are labelled on the page. The other four are method and apply to whatever you are calling. A new account starts with five free requests and no card if you want to follow along.",{"title":64,"description":65},"Can I just scrape it myself in Python?","Often yes, and lesson one is honest about when that is the right call. The threshold is not technical skill — it is how many distinct sites you need and how much you mind being the person who gets paged when one of them redesigns. One site, one developer, no deadline: write it yourself.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":169,"next_step":174},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.","12 min",true,{"title":75,"description":76,"keywords":77},"Declaring an Extraction Schema: Why Your API Fields Are Empty","How to declare the fields a product data API should return, why types and required lists matter, and the common mistake that makes a column come back null on every row.","extraction schema, json schema api, product data fields, api returns null, scraper schema design, structured data extraction",[79,87,92,99,105,109,119,125,154,164],{"type":80,"variant":81,"title":82,"text":83,"cta":84},"callout","product","This lesson uses Scrapewise","The principle is general — every extraction API has some version of a declared schema — but the specific behaviour described here, particularly the required-list rule, is ours. If you are using something else, read it as a prompt to go and test the same question against your own provider, because the answer is rarely in the documentation.",{"label":85,"url":86},"Start with 5 free requests","https://portal.scrapewise.ai/register",{"type":88,"paragraphs":89},"prose",[90,91],"A modern extraction API does not have a fixed response shape. You tell it what a product looks like to you, and it goes and finds those things on the page. That is enormously more flexible than a fixed endpoint, and it moves a responsibility onto you that fixed endpoints never did: if you describe the fields badly, you get bad fields, and nothing anywhere will tell you that is what happened.","The schema is a JSON Schema document. Each field has a name, a type, and usually a description. The names are yours — call the price field price or unit_cost or whatever your downstream job already expects, because renaming it later means touching every consumer.",{"type":93,"title":94,"intro":95,"language":96,"code":97,"caption":98},"code","A minimal product schema","Six fields, which is more than most people need and fewer than most people declare.","json","{\n  \"type\": \"object\",\n  \"properties\": {\n    \"product_title\": { \"type\": \"string\" },\n    \"price\":         { \"type\": [\"number\", \"null\"] },\n    \"currency\":      { \"type\": [\"string\", \"null\"] },\n    \"availability\":  { \"type\": [\"string\", \"null\"] },\n    \"brand\":         { \"type\": [\"string\", \"null\"] },\n    \"sku\":           { \"type\": [\"string\", \"null\"] }\n  },\n  \"required\": [\n    \"product_title\", \"price\", \"currency\",\n    \"availability\", \"brand\", \"sku\"\n  ]\n}","Note that every property is also in required, including the ones that are allowed to be null. That is not a contradiction — see below.",{"type":88,"title":100,"paragraphs":101},"The required list is what makes a field appear",[102,103,104],"This is the single most expensive thing to learn by accident, so here it is plainly: listing a field in required is what makes the extractor emit it. A property that is declared but not required can be omitted entirely, and when it is omitted you do not get a null — you get no key at all, on every row, forever.","The intuition most developers bring is the validation intuition: required means this must not be missing, so leave optional fields out and let them be absent when the page does not have them. That reasoning produces a feed with four columns when you asked for nine, and the symptom looks exactly like \"the site does not publish brand\" rather than \"we never asked for brand\".","The fix is to put every field in required and allow null in its type. You are then saying two different things in the two places: the type says this value may legitimately be absent, and the required list says the key must be present regardless. That is the combination that gives you a stable column set, which is what a warehouse table needs.",{"type":80,"variant":106,"title":107,"text":108},"warning","Field descriptions are not instructions","It is tempting to write a long description explaining what you want — \"the regular price before any discount, excluding delivery\". Test whether that actually changes anything before you rely on it. On our extractor it does not; the required list drives emission and the description is documentation for humans. Writing careful prose into a field nobody reads is a satisfying way to spend an afternoon and change nothing.",{"type":110,"title":111,"intro":112,"items":113},"list","Type choices that save a conversion later","Small decisions, each of which costs an hour downstream if you get it wrong.",[114,115,116,117,118],"Numbers should be numbers, with no currency symbol and no thousands separator. A price that arrives as the string \"1.299,00 €\" is now your parsing problem, and the comma means different things in different markets.","Use [\"number\", \"null\"] rather than [\"string\", \"null\"] for anything you will do arithmetic on, even if the page shows it with a unit.","Keep currency as its own field rather than inferring it from the domain. Multi-market storefronts exist and they will catch you out.","Availability is better as the raw string the page published than as a boolean you guessed. \"Ships in 2–3 weeks\" is not in stock and it is not out of stock.","Always carry the source URL on the row. Every single data question you will be asked begins with \"where did this come from\".",{"type":88,"title":120,"paragraphs":121},"Declare fewer fields than you think you need",[122,123,124],"There is a real cost to a wide schema. More fields means more for the extractor to look for, which means more latency and, on some pricing models, more cost per page. It also means more columns that can be empty, and an empty column is not neutral — it is something a colleague will eventually interpret as a zero.","A better pattern is to scrape only the raw values and compute anything derived afterwards. Unit price, price in a common currency, discount percentage: none of those should be fields the extractor hunts for. They are arithmetic on fields you already have, and doing them after collection means a currency rate change does not require a re-run.","There is a mechanical reason for this too. If a derived value is declared as a scraped field, a post-processing rule that wants to write to that same name usually cannot — the scraped column occupies it. So the derived column stays empty and the rule looks broken when actually it was blocked.",{"type":126,"title":127,"intro":128,"headers":129,"rows":133},"table","Worked example: the same schema, two required arrays","This is the single behaviour that catches most people, and it is worth seeing side by side. Both schemas declare the same five properties. The only difference is what is listed as required — and that difference decides whether a column exists at all. The left-hand column is how a spreadsheet ends up with a header row that moves between runs; the right-hand one is a stable column set with honest empties.",[130,131,132],"Field","required is [\"name\", \"price\"]","required is all five, types allow null",[134,137,140,144,146,150],[135,136,136],"name","\"Acme Widget 500g\"",[138,139,139],"price","19.99",[141,142,143],"sale_price, no promotion on the page","key absent from the row entirely","null",[145,142,143],"pack_size, not printed on the page",[147,148,149],"currency, present on the page","sometimes there, sometimes not","\"EUR\"",[151,152,153],"Columns in the resulting table","varies row by row","five, every row, every run",{"type":110,"title":155,"intro":156,"items":157},"What usually goes wrong","Schema mistakes do not raise errors. They produce a table that is a little bit wrong in a way that is hard to see.",[158,159,160,161,162,163],"Declaring a field and leaving it out of the required array. It is described, it is understood, and it is not emitted — the quietest failure in this whole course.","Writing instructions into the field description. A description says what the field is, not what to do with it; arithmetic, conversions and clean-up all belong downstream.","Declaring a field whose name collides with something computed later. A scraped unit_price column blocks the rule meant to produce one, and the rule then simply never runs.","Asking for thirty fields on the first pass. Each one is another thing that can be wrong, and twenty-five of them will never be read by anybody.","Typing a number as a string because the page shows \"500 g\". Keep the raw text in its own field and let the number field be a number or null, otherwise you are parsing on every read forever.","Scaling to ten thousand pages without reading one complete row. A missing column is obvious in a single row and invisible in a summary count.",{"type":88,"title":165,"paragraphs":166},"Verify with one page before you run ten thousand",[167,168],"Point the configuration at a single known URL and look at the row. Not the row count — the row. Check that every key you declared is present, that the price is a number, that the currency is what you expected, and that the title is the product rather than the category.","If a column is missing, the schema is the first place to look, not the site. If a column is present but null on a page where you can see the value with your own eyes, that is a genuine extraction problem and worth reporting. Those two failures look identical in a dashboard and have completely different fixes.",[170,171,172,173],"Listing a field in the required array is what makes the extractor emit it. Declared-but-not-required fields come back as no key at all.","Put every field in required and allow null in its type. That gives a stable column set with honest empties.","Scrape raw values only. Compute unit prices, conversions and discounts afterwards, or the derived column gets blocked by the scraped one.","Verify against one page and read the whole row before scaling to thousands.",{"text":175,"label":176,"url":177},"Next: why the call that starts a run does not return your data, and how to work with that properly.","Lesson 4: asynchronous runs and polling","/learn/product-data-api/asynchronous-runs-and-polling",2,[180,187,192,193,199,204],{"slug":181,"navTitle":182,"title":183,"summary":184,"time":185,"needsAccount":186},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.","10 min",false,{"slug":188,"navTitle":189,"title":190,"summary":191,"time":185,"needsAccount":186},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":194,"navTitle":195,"title":196,"summary":197,"time":198,"needsAccount":186},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.","11 min",{"slug":200,"navTitle":201,"title":202,"summary":203,"time":198,"needsAccount":186},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"slug":205,"navTitle":206,"title":207,"summary":208,"time":72,"needsAccount":73},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",{"slug":188,"navTitle":189,"title":190,"summary":191,"time":185,"needsAccount":186},{"slug":194,"navTitle":195,"title":196,"summary":197,"time":198,"needsAccount":186},1791047867147]