[{"data":1,"prerenderedAt":235},["ShallowReactive",2],{"learn-lesson-product-data-api-put-the-feed-into-your-stack":3},{"course":4,"lesson":67,"index":202,"outline":203,"prev":233,"next":234},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"product-data-api",3,"Comfortable with HTTP and JSON","6 lessons, about 70 minutes","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Product Data API: A Free 6-Lesson Course for Developers","How to pull product, price and stock data over an API. Authentication, declaring fields, asynchronous runs, pagination, retries without double billing, and putting the feed into your stack. Free, ungated.","product data api, price api, ecommerce api, web scraping api, api key authentication, api documentation, scraper api tutorial, rest api product data","A free developer course on pulling product data over an API","Six written lessons: when an API beats a scraper, auth, declaring fields, async runs, error handling, and shipping the feed into your stack.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course three","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.",{"label":21,"url":22},"Start with lesson one","/learn/product-data-api/when-an-api-beats-a-scraper",{"label":24,"url":25},"See what a run returns","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Decide between an official retailer API, a web data API and writing your own scraper, with reasons you can defend in review","Make an authenticated request and read the response without guessing what the fields mean","Declare a schema so the extractor returns the fields you need rather than the ones it felt like returning","Work with asynchronous runs: start, poll, and handle a job that finishes half-done","Retry a failed request without being billed twice for the same page","Land the result in a warehouse table that does not quietly drift out of date",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Developers who have been told to \"get the competitor prices\" and have discovered the retailer has no public API","Data engineers deciding whether to own the collection layer or buy it","Backend teams wiring a third-party feed into an existing pipeline and wanting to know where it will break","Not written for",[44,45],"Readers looking for a no-code setup — course one covers the same ground without a terminal","Anyone wanting a specific vendor's endpoint reference. This is the shape of the problem; the reference lives in that vendor's docs.",{"title":47,"intro":48},"The six lessons","Lessons one and two are short and assume nothing but curl. From three onwards you will get more out of it with a key in your hand.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions developers ask in the first ten minutes.",[54,57,60,63],{"title":55,"description":56},"Does this teach a specific vendor's API?","No, deliberately. Endpoint names change and a course that hardcodes them rots. What does not change is the shape: authenticate, declare what you want, start a run, poll it, handle partial results, retry safely. Learn the shape and any vendor's reference becomes a lookup rather than a tutorial.",{"title":58,"description":59},"Why is it not just a GET that returns the price?","Because fetching a page takes seconds and fetching ten thousand takes hours. Any API that hides that behind a synchronous request is either very slow or quietly returning you a cached number. Lesson four is about why the asynchronous shape exists and how to work with it instead of around it.",{"title":61,"description":62},"Do I need a Scrapewise account?","Lessons three and six include a Scrapewise-specific walkthrough and are labelled on the page. The other four are method and apply to whatever you are calling. A new account starts with five free requests and no card if you want to follow along.",{"title":64,"description":65},"Can I just scrape it myself in Python?","Often yes, and lesson one is honest about when that is the right call. The threshold is not technical skill — it is how many distinct sites you need and how much you mind being the person who gets paged when one of them redesigns. One site, one developer, no deadline: write it yourself.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":193,"next_step":198},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.","12 min",true,{"title":75,"description":76,"keywords":77},"Loading a Product Price Feed Into a Warehouse Without Drift","How to schedule collection, design the landing table, keep price history, and detect the day a feed silently halves. Practical patterns for a durable data pipeline.","price feed warehouse, product data pipeline, data loading schema, price history table, feed monitoring, etl product data",[79,87,92,98,119,125,135,139,145,178,188],{"type":80,"variant":81,"title":82,"text":83,"cta":84},"callout","product","This lesson uses Scrapewise","The scheduling and export mechanics described are ours. The table design and the monitoring are not vendor-specific and are the part worth copying regardless of where your data comes from.",{"label":85,"url":86},"Start with 5 free requests","https://portal.scrapewise.ai/register",{"type":88,"paragraphs":89},"prose",[90,91],"The integration is working, the rows come back, and now the real question: where do they go, and will the answer still be right in six months?","Most feeds do not die dramatically. They drift. A column that used to be populated goes mostly null and nobody notices because nobody queries it. A site is dropped from the schedule during a cleanup and the dashboard quietly covers one fewer competitor. The table has a last_updated column that stopped updating. All three are invisible if the only thing you check is whether the pipeline ran.",{"type":88,"title":93,"paragraphs":94},"Decide the cadence from the decision, not the data",[95,96,97],"The common instinct is to collect as often as the provider allows. That is backwards. The right frequency is set by how often somebody acts on the number.","If your repricing job runs once a day at six in the morning, collecting hourly buys you nothing except twenty-three discarded datasets and a bill. If a buyer checks a dashboard on Monday mornings, daily collection is already generous. Hourly makes sense for a narrow set of volatile, high-value SKUs, and almost never for the whole catalogue.","Set the collection to finish comfortably before the thing that consumes it starts. Comfortably means with enough margin that a slow run does not mean a stale decision — and a slow run is normal, because run duration is a property of how the sites behaved that day.",{"type":99,"title":100,"intro":101,"headers":102,"rows":106},"table","Three ways to get the rows out, and when each fits","Most platforms offer all three. They are not interchangeable.",[103,104,105],"Method","Good for","The catch",[107,111,115],[108,109,110],"Scheduled export to a file","Warehouse loads, BI tools, anything batch","You need a loader on the other end, and somebody has to notice when a file does not arrive",[112,113,114],"Pull the results over the API","Full control, custom transforms, backfills","Pagination and retry logic are now yours, and so is the scheduling",[116,117,118],"Webhook on completion","Triggering a downstream job the moment data is ready","Deliveries get lost. Pair with a poll, as in lesson four.",{"type":88,"title":120,"paragraphs":121},"Land it raw, then transform",[122,123,124],"Write what the API returned into a landing table, unmodified, with a collection timestamp and the run id. Then build your cleaned view on top of it. This is standard advice and it is standard for a reason: the day somebody asks why Tuesday's price looks wrong, you want to be able to see what actually arrived on Tuesday rather than what your transform made of it.","It also makes a transform bug recoverable. If the cleaned table is the only copy and your parser mangled a currency for a fortnight, that fortnight is gone. If the landing table is intact you re-run the transform and the fortnight comes back.","Keep the source URL on every row through every layer. It is the only thing that lets a human verify a number by opening a page, and it will be asked for.",{"type":126,"title":127,"intro":128,"items":129},"list","Columns worth having that people leave out","Each of these exists to answer a question somebody will eventually ask.",[130,131,132,133,134],"collected_at — when the value was read, not when the row was written. These differ and the difference is the age of your data.","run_id — so a suspicious batch can be traced to one run and compared against its neighbours","source_url — the exact page, so a number can be verified by a human in one click","raw_price_text — what the page actually said, before parsing. Invaluable the first time a decimal separator goes wrong.","A row per observation rather than an updated row per product, so that price history exists by construction rather than being added later at great expense",{"type":80,"variant":136,"title":137,"text":138},"warning","Overwriting is how you lose history you did not know you needed","An update-in-place table answers \"what is the price now\" and nothing else. The first interesting question anyone asks of a price feed is \"when did they change it\", and if you have been overwriting, the honest answer is that you do not know and cannot find out. Append-only costs more storage than you will ever notice and it is the single decision that most determines whether this dataset is useful in a year.",{"type":88,"title":140,"paragraphs":141},"Monitor the shape of the data, not the health of the job",[142,143,144],"A green pipeline run tells you the code did not throw. It says nothing about whether the data is right, and the failures that matter are almost all of the second kind.","Three checks catch most of it. Row count against the trailing average, per site, with an alert on a meaningful drop — a feed halving is a far more common failure than a feed stopping. Null rate per column against its own baseline, because a column going empty is a redesign signal. And a freshness check that actually compares collected_at to now, rather than trusting that the scheduler ran.","The reason to do this per site rather than in aggregate is that aggregate numbers hide exactly the failures you care about. Forty sites holding steady and one going to zero is a two and a half per cent drop in the total, which no sensible threshold will catch, and it is also one competitor having completely vanished from your pricing decisions.",{"type":99,"title":146,"intro":147,"headers":148,"rows":152},"Worked example: what update-in-place destroys","Four observations of one SKU, stored two ways. The left-hand table is what almost everyone builds first, because it matches how a price feels — a current value. The right-hand one answers the question that always gets asked. The cost of the right-hand version is 365 rows a year per SKU per competitor, which for 1,200 SKUs and five competitors is about 2.2 million rows: unremarkable for any database built this decade, and more than any spreadsheet should be asked to hold.",[149,150,151],"Observation","Update-in-place","Append-only",[153,157,160,164,167,171,174],[154,155,156],"Monday, €229.00","one row: €229.00","row 1",[158,155,159],"Tuesday, €229.00","row 2",[161,162,163],"Wednesday, €189.00","one row: €189.00","row 3",[165,162,166],"Thursday, €189.00","row 4",[168,169,170],"\"When did they drop it, and from what?\"","unanswerable","Wednesday, from €229.00",[172,169,173],"\"How long has the promotion run?\"","two days so far",[175,176,177],"Rows after a year of daily runs, one SKU","1","365",{"type":126,"title":179,"intro":180,"items":181},"What usually goes wrong","Six months in, a feed is either still trusted or quietly worked around. These are the decisions that determine which.",[182,183,184,185,186,187],"Updating in place. The cheapest decision on day one and the most regretted, because the history cannot be reconstructed afterwards.","Transforming before landing. If the raw response is discarded, a bug in the transform costs you the data instead of an afternoon of reprocessing.","Setting cadence from what the provider allows. Set it from how often somebody actually changes a price; everything above that is pages you pay for and nobody reads.","Monitoring the job instead of the shape. Row count and null rate per site against a trailing baseline catch the degradations. A green job catches almost nothing.","Monitoring in aggregate. One competitor out of five dropping to zero moves the total by a fifth, which sits comfortably under any threshold you would have set.","Leaving the coverage gaps undocumented. The feed does not include third-party marketplace sellers, or a market nobody configured, or anything behind a login — and the first person surprised by that will be the one presenting from it.",{"type":88,"title":189,"paragraphs":190},"Write down what the feed does not cover",[191,192],"Every feed has holes: sites that block, pages that do not publish a price, categories that were never in scope. These are fine. What is not fine is that they live in one engineer's head.","Keep a short document next to the pipeline that lists which sites are in scope, which are excluded and why, and which fields are known to be unreliable where. It takes twenty minutes to write and it prevents the single most damaging thing a data feed can do, which is to be quietly interpreted as complete by somebody who was not there when it was built.",[194,195,196,197],"Set collection cadence from how often someone acts on the number, not from what the provider allows.","Land the raw response first, transform on top. A transform bug is then recoverable; otherwise the data is gone.","Append a row per observation. Update-in-place destroys the price history that is the first thing anyone asks for.","Monitor row count and null rate per site against a baseline. Aggregate monitoring hides the one competitor that vanished.",{"text":199,"label":200,"url":201},"That is the course. The next one is about the thing that breaks this feed most often: the sites themselves changing, and occasionally deciding they would rather not be read.","Course four: keeping scrapers alive","/learn/keep-scrapers-alive",5,[204,211,216,221,227,232],{"slug":205,"navTitle":206,"title":207,"summary":208,"time":209,"needsAccount":210},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.","10 min",false,{"slug":212,"navTitle":213,"title":214,"summary":215,"time":209,"needsAccount":210},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":217,"navTitle":218,"title":219,"summary":220,"time":72,"needsAccount":73},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.",{"slug":222,"navTitle":223,"title":224,"summary":225,"time":226,"needsAccount":210},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.","11 min",{"slug":228,"navTitle":229,"title":230,"summary":231,"time":226,"needsAccount":210},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":228,"navTitle":229,"title":230,"summary":231,"time":226,"needsAccount":210},null,1791047867187]