[{"data":1,"prerenderedAt":238},["ShallowReactive",2],{"learn-lesson-product-data-api-asynchronous-runs-and-polling":3},{"course":4,"lesson":67,"index":6,"outline":206,"prev":236,"next":237},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"product-data-api",3,"Comfortable with HTTP and JSON","6 lessons, about 70 minutes","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Product Data API: A Free 6-Lesson Course for Developers","How to pull product, price and stock data over an API. Authentication, declaring fields, asynchronous runs, pagination, retries without double billing, and putting the feed into your stack. Free, ungated.","product data api, price api, ecommerce api, web scraping api, api key authentication, api documentation, scraper api tutorial, rest api product data","A free developer course on pulling product data over an API","Six written lessons: when an API beats a scraper, auth, declaring fields, async runs, error handling, and shipping the feed into your stack.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course three","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.",{"label":21,"url":22},"Start with lesson one","/learn/product-data-api/when-an-api-beats-a-scraper",{"label":24,"url":25},"See what a run returns","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Decide between an official retailer API, a web data API and writing your own scraper, with reasons you can defend in review","Make an authenticated request and read the response without guessing what the fields mean","Declare a schema so the extractor returns the fields you need rather than the ones it felt like returning","Work with asynchronous runs: start, poll, and handle a job that finishes half-done","Retry a failed request without being billed twice for the same page","Land the result in a warehouse table that does not quietly drift out of date",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Developers who have been told to \"get the competitor prices\" and have discovered the retailer has no public API","Data engineers deciding whether to own the collection layer or buy it","Backend teams wiring a third-party feed into an existing pipeline and wanting to know where it will break","Not written for",[44,45],"Readers looking for a no-code setup — course one covers the same ground without a terminal","Anyone wanting a specific vendor's endpoint reference. This is the shape of the problem; the reference lives in that vendor's docs.",{"title":47,"intro":48},"The six lessons","Lessons one and two are short and assume nothing but curl. From three onwards you will get more out of it with a key in your hand.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions developers ask in the first ten minutes.",[54,57,60,63],{"title":55,"description":56},"Does this teach a specific vendor's API?","No, deliberately. Endpoint names change and a course that hardcodes them rots. What does not change is the shape: authenticate, declare what you want, start a run, poll it, handle partial results, retry safely. Learn the shape and any vendor's reference becomes a lookup rather than a tutorial.",{"title":58,"description":59},"Why is it not just a GET that returns the price?","Because fetching a page takes seconds and fetching ten thousand takes hours. Any API that hides that behind a synchronous request is either very slow or quietly returning you a cached number. Lesson four is about why the asynchronous shape exists and how to work with it instead of around it.",{"title":61,"description":62},"Do I need a Scrapewise account?","Lessons three and six include a Scrapewise-specific walkthrough and are labelled on the page. The other four are method and apply to whatever you are calling. A new account starts with five free requests and no card if you want to follow along.",{"title":64,"description":65},"Can I just scrape it myself in Python?","Often yes, and lesson one is honest about when that is the right call. The threshold is not technical skill — it is how many distinct sites you need and how much you mind being the person who gets paged when one of them redesigns. One site, one developer, no deadline: write it yourself.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":197,"next_step":202},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.","11 min",false,{"title":75,"description":76,"keywords":77},"Asynchronous API Runs: Polling, Job State and Partial Results","Why web data APIs return a job id instead of data, how to poll correctly with backoff, how to handle pagination, and what to do when a run completes partially.","asynchronous api, polling api job status, api pagination, long running request, webhook vs polling, partial results",[79,85,99,105,110,116,122,147,186],{"type":80,"paragraphs":81},"prose",[82,83,84],"The first thing that surprises people coming from ordinary REST is that starting a collection does not give you any data. It gives you an identifier. The work happens somewhere else, over minutes or hours, and you come back for the result.","This is not an architectural preference. Fetching a single page takes seconds in the best case, and much longer when the site is slow, when a retry is needed, or when a headless browser has to render it. A synchronous endpoint over ten thousand URLs would hold a connection open for hours, and the first network blip would lose the entire result. So the API hands you a job.","The consequence for your code is that there are now three distinct operations where you expected one: start, check, collect.",{"type":86,"title":87,"intro":88,"items":89},"steps","The three-step shape","Every asynchronous data API is some variation of this, whatever the endpoints are called.",[90,93,96],{"title":91,"text":92},"Start the run","Post the configuration, or the identifier of a saved configuration, and get back a job or run id. Store that id somewhere durable immediately. If your process dies here and the id was only in memory, you have paid for a run you cannot collect.",{"title":94,"text":95},"Poll until it settles","Ask for the run's state on an interval until it reaches a terminal state. Terminal usually means finished, failed or stopped. Anything else means keep waiting.",{"title":97,"text":98},"Collect the rows","Fetch the results, usually paginated. This is a separate call and often a separate rate limit, which matters because collecting a large run can take longer than you expect.",{"type":80,"title":100,"paragraphs":101},"How to poll without being rude",[102,103,104],"A fixed one-second poll on a job that takes forty minutes is two and a half thousand pointless requests, and on some providers it is two and a half thousand requests against your rate limit.","Use exponential backoff with a ceiling. Start at a few seconds, double each time, cap at thirty or sixty seconds. For a job you expect to run for an hour you will make a few dozen calls rather than thousands, and you will still notice completion within a minute of it happening.","If the provider offers a webhook, use it and keep the polling as a fallback rather than deleting it. Webhooks get lost — a deploy restarts your process mid-delivery, a firewall rule changes, the retry policy is less generous than you assumed. A slow poll behind a fast webhook costs almost nothing and converts a silent stall into a late notification.",{"type":106,"variant":107,"title":108,"text":109},"callout","warning","Set a deadline, and decide now what happens when it passes","A run that never reaches a terminal state is the failure mode that hurts most, because nothing errors. Your poller just keeps politely asking. Pick a maximum wall-clock time based on the size of the run, and when it elapses, stop the run explicitly rather than abandoning it — an abandoned run on most platforms carries on working and carries on billing. The decision to make in advance is whether a timed-out run's partial rows are usable or discarded, because making that call at three in the morning goes badly.",{"type":80,"title":111,"paragraphs":112},"Pagination and the row cap nobody mentions",[113,114,115],"Results come back in pages. The mechanics are ordinary — a limit, an offset or a cursor, a loop — but there are two traps.","The first is that sample and preview endpoints frequently cap out at a round number like a hundred rows. If you test with one of those and then build your pagination against it, you will conclude the loop works and that runs are smaller than they are. Check whether the endpoint you are looping over is the full-results endpoint or the preview one.","The second is that the count you use to decide when to stop should come from the run's own metadata, not from counting what you have received. If a page comes back short because of a transient error and your loop uses received-count as its terminator, you will stop early and treat a partial collection as a complete one.",{"type":80,"title":117,"paragraphs":118},"Partial results are the normal case",[119,120,121],"A run over several thousand URLs will not return several thousand rows. Some pages will be gone, some will be a redirect to a category listing, some will be behind a bot wall on the day you ran it. A ninety per cent return is a good run.","So completion and completeness are different questions, and your code should ask both. The run finished — fine. Did it return roughly what the last one did? That second question is the one that catches real problems, and lesson five of course one covers how to baseline it.","The thing not to do is treat a short run as a failure and retry the whole thing. You will pay twice to collect the same ninety per cent, and the ten per cent that failed will mostly fail again, because the reason was the page and not the attempt.",{"type":123,"title":124,"intro":125,"headers":126,"rows":130},"table","Terminal states and what each one means for your data","Names vary. The categories do not.",[127,128,129],"State","Meaning","Are the rows usable?",[131,135,139,143],[132,133,134],"Finished","The run processed its whole input list","Yes, subject to the usual per-page failures",[136,137,138],"Stopped","Something ended it early — you, a cap, or a deadline","Usually yes, but the input was not fully covered",[140,141,142],"Failed","The run itself could not proceed, often a config error","Treat as none; fix the configuration first",[144,145,146],"Finished with errors","Completed, but a meaningful share of pages did not return","Yes, and the error breakdown is the thing to read",{"type":123,"title":148,"intro":149,"headers":150,"rows":155},"Worked example: a polling schedule that does not hammer","Exponential backoff with a thirty-second cap, against a run that takes four minutes. Twelve requests instead of the 270 a one-second poll would have made, and the longest you ever sit on a finished run is thirty seconds.",[151,152,153,154],"Poll","Wait before it","Elapsed","Status",[156,161,165,169,173,177,181],[157,158,159,160],"1","2s","0:02","running",[162,163,164,160],"2","4s","0:06",[166,167,168,160],"3","8s","0:14",[170,171,172,160],"4","16s","0:30",[174,175,176,160],"5","30s (cap reached)","1:00",[178,179,180,160],"6 to 11","30s each","4:00",[182,183,184,185],"12","30s","4:30","succeeded",{"type":187,"title":188,"intro":189,"items":190},"list","What usually goes wrong","Asynchronous collection fails in ways that look like success, which is why the counts matter more than the state.",[191,192,193,194,195,196],"Polling on a fixed one-second interval. That is 270 requests where twelve would do, and on a provider with a global rate limit the poller can starve the collection it is waiting for.","Not storing the run id before the first poll. The process restarts, the id is gone, and the run carries on — and carries on billing — with nobody collecting it.","No wall-clock deadline. A stalled run polls forever, and forever is a line on an invoice.","Treating finished as complete. A run can terminate successfully having fetched 900 of 1,200 pages. Ask for the state and the counts, and compare the counts against the last good run.","Ignoring the row cap. Many endpoints return at most a few hundred rows per page, so a run that produced 4,000 rows hands back 100 and a cursor, and a naive client records 100.","Throwing away partial results on failure. Eighty per cent of a run is usually worth keeping, provided the gap is recorded rather than quietly absorbed into an average.",[198,199,200,201],"Start, poll, collect. Store the run id durably the moment you get it.","Poll with exponential backoff and a cap. Use webhooks if offered, but keep a slow poll as a fallback.","Set a wall-clock deadline and stop a stalled run explicitly — abandoned runs keep billing.","Partial results are normal. Completion and completeness are two different questions and you should ask both.",{"text":203,"label":204,"url":205},"Next: retrying properly, including how to avoid paying twice for the same page.","Lesson 5: errors, retries and double billing","/learn/product-data-api/errors-retries-and-double-billing",[207,213,218,225,226,231],{"slug":208,"navTitle":209,"title":210,"summary":211,"time":212,"needsAccount":73},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.","10 min",{"slug":214,"navTitle":215,"title":216,"summary":217,"time":212,"needsAccount":73},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":219,"navTitle":220,"title":221,"summary":222,"time":223,"needsAccount":224},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.","12 min",true,{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":227,"navTitle":228,"title":229,"summary":230,"time":72,"needsAccount":73},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"slug":232,"navTitle":233,"title":234,"summary":235,"time":223,"needsAccount":224},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",{"slug":219,"navTitle":220,"title":221,"summary":222,"time":223,"needsAccount":224},{"slug":227,"navTitle":228,"title":229,"summary":230,"time":72,"needsAccount":73},1791047867158]