[{"data":1,"prerenderedAt":236},["ShallowReactive",2],{"learn-lesson-product-data-api-errors-retries-and-double-billing":3},{"course":4,"lesson":67,"index":203,"outline":204,"prev":234,"next":235},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"product-data-api",3,"Comfortable with HTTP and JSON","6 lessons, about 70 minutes","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Product Data API: A Free 6-Lesson Course for Developers","How to pull product, price and stock data over an API. Authentication, declaring fields, asynchronous runs, pagination, retries without double billing, and putting the feed into your stack. Free, ungated.","product data api, price api, ecommerce api, web scraping api, api key authentication, api documentation, scraper api tutorial, rest api product data","A free developer course on pulling product data over an API","Six written lessons: when an API beats a scraper, auth, declaring fields, async runs, error handling, and shipping the feed into your stack.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course three","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.",{"label":21,"url":22},"Start with lesson one","/learn/product-data-api/when-an-api-beats-a-scraper",{"label":24,"url":25},"See what a run returns","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Decide between an official retailer API, a web data API and writing your own scraper, with reasons you can defend in review","Make an authenticated request and read the response without guessing what the fields mean","Declare a schema so the extractor returns the fields you need rather than the ones it felt like returning","Work with asynchronous runs: start, poll, and handle a job that finishes half-done","Retry a failed request without being billed twice for the same page","Land the result in a warehouse table that does not quietly drift out of date",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Developers who have been told to \"get the competitor prices\" and have discovered the retailer has no public API","Data engineers deciding whether to own the collection layer or buy it","Backend teams wiring a third-party feed into an existing pipeline and wanting to know where it will break","Not written for",[44,45],"Readers looking for a no-code setup — course one covers the same ground without a terminal","Anyone wanting a specific vendor's endpoint reference. This is the shape of the problem; the reference lives in that vendor's docs.",{"title":47,"intro":48},"The six lessons","Lessons one and two are short and assume nothing but curl. From three onwards you will get more out of it with a key in your hand.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions developers ask in the first ten minutes.",[54,57,60,63],{"title":55,"description":56},"Does this teach a specific vendor's API?","No, deliberately. Endpoint names change and a course that hardcodes them rots. What does not change is the shape: authenticate, declare what you want, start a run, poll it, handle partial results, retry safely. Learn the shape and any vendor's reference becomes a lookup rather than a tutorial.",{"title":58,"description":59},"Why is it not just a GET that returns the price?","Because fetching a page takes seconds and fetching ten thousand takes hours. Any API that hides that behind a synchronous request is either very slow or quietly returning you a cached number. Lesson four is about why the asynchronous shape exists and how to work with it instead of around it.",{"title":61,"description":62},"Do I need a Scrapewise account?","Lessons three and six include a Scrapewise-specific walkthrough and are labelled on the page. The other four are method and apply to whatever you are calling. A new account starts with five free requests and no card if you want to follow along.",{"title":64,"description":65},"Can I just scrape it myself in Python?","Often yes, and lesson one is honest about when that is the right call. The threshold is not technical skill — it is how many distinct sites you need and how much you mind being the person who gets paged when one of them redesigns. One site, one developer, no deadline: write it yourself.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":194,"next_step":199},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.","11 min",false,{"title":75,"description":76,"keywords":77},"API Retries Without Double Billing: Idempotency and Backoff","Which API errors to retry and which to stop on, how idempotency keys prevent duplicate charges, exponential backoff with jitter, and rate limit handling.","api retry strategy, idempotency key, exponential backoff, rate limit 429, duplicate api charge, api error handling",[79,84,118,123,130,136,141,151,178,188],{"type":80,"paragraphs":81},"prose",[82,83],"Retry logic is where a tidy integration becomes an expensive one. The default instinct — wrap the call in a loop, try three times, move on — is wrong in two directions at once. It retries things that will never succeed, and it retries things that already succeeded.","The second of those is the one that costs money. A request can time out on your side after the server has already done the work. From your perspective nothing happened; from the provider's, a run started and a bill was incurred. Retry and you have two.",{"type":85,"title":86,"intro":87,"headers":88,"rows":92},"table","Retry or stop","The useful split is not by status code but by whether trying again could plausibly change the outcome.",[89,90,91],"Failure","Retry?","Why",[93,97,101,105,109,112,115],[94,95,96],"Connection reset, DNS blip, timeout","Yes, with backoff","Genuinely transient. The next attempt is a different network moment.",[98,99,100],"500, 502, 503, 504","Yes, with backoff and a cap","Their side, usually brief. Cap it so a long outage does not become an infinite loop.",[102,103,104],"429","Yes, but slower","Honour Retry-After. Also reduce concurrency — retrying at the same rate just re-earns the 429.",[106,107,108],"400, 422","No","The payload is wrong. It will be wrong on the next attempt too.",[110,107,111],"401, 403","The key is missing, revoked or unscoped. Alert a human instead.",[113,107,114],"404","Check the path against the reference. Retrying a typo is just a slower typo.",[116,107,117],"200 with an error body","Deterministic. Read the body and handle it as the failure it is.",{"type":80,"title":119,"paragraphs":120},"Backoff needs jitter",[121,122],"Exponential backoff alone has a well-known failure mode. If fifty of your workers all fail at the same moment — which is exactly what happens during a brief provider outage — they all back off by the same amount and all retry at the same instant. You have rebuilt the thundering herd with extra steps, and your retry storm is what keeps the recovery from happening.","Add randomness. Wait somewhere between zero and the backoff interval rather than exactly the interval. It is a one-line change and it is the difference between a recovery and a second outage.",{"type":124,"title":125,"intro":126,"language":127,"code":128,"caption":129},"code","Backoff with jitter, in the smallest honest form","Deliberately boring. The important parts are the cap, the jitter, and that it only loops on retryable failures.","python","import random, time\n\ndef with_retries(call, attempts=5, base=2.0, cap=60.0):\n    for n in range(attempts):\n        try:\n            return call()\n        except Retryable:\n            if n == attempts - 1:\n                raise\n            delay = min(cap, base * (2 ** n))\n            time.sleep(random.uniform(0, delay))","Retryable is your own exception, raised only for the rows in the table above that say yes. Everything else should propagate immediately.",{"type":80,"title":131,"paragraphs":132},"Idempotency keys",[133,134,135],"An idempotency key is a value you generate and attach to a request that starts work. If the provider supports it, a second request carrying the same key does not start a second run — it returns the result of the first one.","This is the clean answer to the timed-out-but-succeeded problem. Generate the key before the first attempt, reuse it on every retry of that same logical operation, and a retry becomes safe rather than expensive. Generating a fresh key per attempt defeats the entire mechanism, which is a surprisingly common mistake because the generation often lives inside the retry loop by accident.","Two practical notes. Support is usually opt-in, so it does nothing unless you send the header. And it is often supported on more methods than you would guess, including deletes, which matters because a timed-out delete is just as ambiguous as a timed-out create.",{"type":137,"variant":138,"title":139,"text":140},"callout","warning","Rate limits are frequently global, not per-endpoint","It is natural to assume that a heavy endpoint has its own budget and that a cheap one, like checking a job's status, does not count. Often untrue. If the limit is applied globally, an aggressive poller can consume the allowance your actual collection needs, and the symptom is that the collection starts failing while the polling keeps working perfectly. If you are seeing 429s, count all your calls, not just the obvious ones.",{"type":142,"title":143,"intro":144,"items":145},"list","Log enough to diagnose it later","What you want in the log line at the moment of failure, because none of it is recoverable afterwards.",[146,147,148,149,150],"The run or request id the provider gave you — without it, support cannot help and neither can you","The status code and the first part of the response body, not just the exception message","Which attempt number this was, so a one-off is distinguishable from a systematic failure","The timestamp in UTC, because the provider's logs are in UTC and correlating across zones at 3am is how mistakes happen","A durable note of whether work may have started, so the next operator knows whether a retry is safe",{"type":85,"title":152,"intro":153,"headers":154,"rows":159},"Worked example: how one timeout becomes three invoices","A start-run request times out at the client after thirty seconds. The server received it and is working on it. Without an idempotency key every retry is a new run, and all of them finish. With one key generated before the first attempt, attempts two and three return run A's id and the bill is one run — but the key has to exist before attempt one. A key generated inside the retry loop is a new key each time and does nothing whatsoever.",[155,156,157,158],"Attempt","What the client sees","What the server does","Runs in flight",[160,164,168,173],[161,162,163,161],"1","timeout after 30s","accepted, starts run A",[165,162,166,167],"2, a retry","accepted, starts run B","2",[169,170,171,172],"3, a retry","200, run id C","accepted, starts run C","3",[174,175,176,177],"Outcome","one run id","three complete runs","three times the pages, a third of them collected",{"type":142,"title":179,"intro":180,"items":181},"What usually goes wrong","Retry logic is written on a good day and exercised on a bad one, which is why these are so common.",[182,183,184,185,186,187],"Retrying a 400. Nothing about a second attempt changes a malformed request, and some providers charge for the attempt anyway.","Backoff without jitter. Every client that failed during the outage retries on exactly the same schedule and rebuilds the stampede at the moment the service is trying to recover.","Generating the idempotency key inside the retry loop. A fresh key per attempt is identical to no key, and it reads as correct in the code.","Assuming the rate limit is per endpoint. It is frequently global, so an aggressive poller on one route can 429 the route that actually matters.","Retrying a quota or payment error. That is the class that means stop: the system is working correctly and telling you that you are out of credit.","Logging that a call failed without logging what was sent. A month later you have a count of failures and no way to reproduce a single one.",{"type":80,"title":189,"paragraphs":190},"The failure that means stop, not try harder",[191,192,193],"There is one category where retrying is not just useless but actively the wrong instinct: when the page is refusing you rather than failing.","A bot wall that has decided your traffic is automated does not give you a flaky response that works on the third attempt. It gives you a consistent refusal, and hammering it makes the classification worse. Three retries against a wall is three times the cost for the same answer, plus a stronger signal that you are what the wall thinks you are.","This is the main subject of the next course, and the distinction worth carrying into it is between a page that failed and a page that declined. Your retry logic can only help with the first.",[195,196,197,198],"Split failures by whether a second attempt could plausibly change the outcome, not by status code family.","Exponential backoff without jitter rebuilds the thundering herd during exactly the outage you are trying to survive.","Generate the idempotency key before the first attempt and reuse it on every retry, or it does nothing.","Rate limits are often global. An aggressive poller can starve the collection it is polling for.",{"text":200,"label":201,"url":202},"Next: getting the feed out of the API and into something your team actually queries.","Lesson 6: putting the feed into your stack","/learn/product-data-api/put-the-feed-into-your-stack",4,[205,211,216,223,228,229],{"slug":206,"navTitle":207,"title":208,"summary":209,"time":210,"needsAccount":73},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.","10 min",{"slug":212,"navTitle":213,"title":214,"summary":215,"time":210,"needsAccount":73},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":217,"navTitle":218,"title":219,"summary":220,"time":221,"needsAccount":222},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.","12 min",true,{"slug":224,"navTitle":225,"title":226,"summary":227,"time":72,"needsAccount":73},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":230,"navTitle":231,"title":232,"summary":233,"time":221,"needsAccount":222},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",{"slug":224,"navTitle":225,"title":226,"summary":227,"time":72,"needsAccount":73},{"slug":230,"navTitle":231,"title":232,"summary":233,"time":221,"needsAccount":222},1791047867165]