[{"data":1,"prerenderedAt":222},["ShallowReactive",2],{"learn-lesson-product-data-api-when-an-api-beats-a-scraper":3},{"course":4,"lesson":67,"index":189,"outline":190,"prev":220,"next":221},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"product-data-api",3,"Comfortable with HTTP and JSON","6 lessons, about 70 minutes","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Product Data API: A Free 6-Lesson Course for Developers","How to pull product, price and stock data over an API. Authentication, declaring fields, asynchronous runs, pagination, retries without double billing, and putting the feed into your stack. Free, ungated.","product data api, price api, ecommerce api, web scraping api, api key authentication, api documentation, scraper api tutorial, rest api product data","A free developer course on pulling product data over an API","Six written lessons: when an API beats a scraper, auth, declaring fields, async runs, error handling, and shipping the feed into your stack.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course three","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.",{"label":21,"url":22},"Start with lesson one","/learn/product-data-api/when-an-api-beats-a-scraper",{"label":24,"url":25},"See what a run returns","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Decide between an official retailer API, a web data API and writing your own scraper, with reasons you can defend in review","Make an authenticated request and read the response without guessing what the fields mean","Declare a schema so the extractor returns the fields you need rather than the ones it felt like returning","Work with asynchronous runs: start, poll, and handle a job that finishes half-done","Retry a failed request without being billed twice for the same page","Land the result in a warehouse table that does not quietly drift out of date",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Developers who have been told to \"get the competitor prices\" and have discovered the retailer has no public API","Data engineers deciding whether to own the collection layer or buy it","Backend teams wiring a third-party feed into an existing pipeline and wanting to know where it will break","Not written for",[44,45],"Readers looking for a no-code setup — course one covers the same ground without a terminal","Anyone wanting a specific vendor's endpoint reference. This is the shape of the problem; the reference lives in that vendor's docs.",{"title":47,"intro":48},"The six lessons","Lessons one and two are short and assume nothing but curl. From three onwards you will get more out of it with a key in your hand.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions developers ask in the first ten minutes.",[54,57,60,63],{"title":55,"description":56},"Does this teach a specific vendor's API?","No, deliberately. Endpoint names change and a course that hardcodes them rots. What does not change is the shape: authenticate, declare what you want, start a run, poll it, handle partial results, retry safely. Learn the shape and any vendor's reference becomes a lookup rather than a tutorial.",{"title":58,"description":59},"Why is it not just a GET that returns the price?","Because fetching a page takes seconds and fetching ten thousand takes hours. Any API that hides that behind a synchronous request is either very slow or quietly returning you a cached number. Lesson four is about why the asynchronous shape exists and how to work with it instead of around it.",{"title":61,"description":62},"Do I need a Scrapewise account?","Lessons three and six include a Scrapewise-specific walkthrough and are labelled on the page. The other four are method and apply to whatever you are calling. A new account starts with five free requests and no card if you want to follow along.",{"title":64,"description":65},"Can I just scrape it myself in Python?","Often yes, and lesson one is honest about when that is the right call. The threshold is not technical skill — it is how many distinct sites you need and how much you mind being the person who gets paged when one of them redesigns. One site, one developer, no deadline: write it yourself.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":180,"next_step":185},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.","10 min",false,{"title":75,"description":76,"keywords":77},"Product API vs Web Scraper vs Dataset: How to Choose","Official retailer APIs, web data APIs and self-written scrapers compared on coverage, maintenance and cost. The question that actually decides it.","product api vs scraper, web data api, retailer api access, build or buy scraper, ecommerce data sourcing",[79,84,114,120,125,135,165,174],{"type":80,"paragraphs":81},"prose",[82,83],"If you search for \"Walmart API documentation\" or \"Target API key\" you will find thousands of other people who searched for the same thing. You will also find that the answer is usually some version of \"apply to be a partner\", and that the partner API, when you get it, covers your own listings rather than everybody else's. This is not an oversight. A retailer publishes an API so that suppliers can manage their own catalogue; publishing one that lets a competitor read every price would be a strange thing to do on purpose.","So the real choice is not between an official API and something else. It is between three options, and most teams pick one before they have articulated what they are actually optimising for.",{"type":85,"title":86,"intro":87,"headers":88,"rows":93},"table","The three ways to get product data","Coverage means which sites you can reach. Maintenance means who gets paged when one of them changes.",[89,90,91,92],"Route","Coverage","Maintenance","Best when",[94,99,104,109],[95,96,97,98],"Official retailer API","One retailer, usually only your own listings","Theirs, and it is versioned","You are a seller on that platform managing your own catalogue",[100,101,102,103],"Web data API","Any public page, if the vendor can reach it","Theirs, but you still own the schema","You need several sites and want one integration",[105,106,107,108],"Your own scraper","Whatever you build, exactly","Yours, forever","One or two sites, stable, and you want full control",[110,111,112,113],"A bought dataset","Broad but fixed, and already stale","Nobody's — it is a file","You need a one-off snapshot for analysis, not a feed",{"type":80,"title":115,"paragraphs":116},"The question that actually decides it",[117,118,119],"Not \"can we build this\". You can. A competent developer can read a price off a page in an afternoon, and the first version will work.","The question is: how many distinct sites, and how long does this have to keep working? Those two numbers together are the whole decision. One site for a quarter is a script. Forty sites indefinitely is an operations problem that happens to involve HTTP, and the cost is not the code — it is the on-call rotation, the day a redesign breaks nine of them at once, and the proxy bill you did not budget for.","The honest crossover, from running both: somewhere around three to five sites, or the moment the feed becomes something a person makes a pricing decision from. Below that, write it. Above that, the integration you maintain should be one API rather than forty parsers.",{"type":121,"variant":122,"title":123,"text":124},"callout","warning","The cost nobody puts in the build-versus-buy spreadsheet","Build estimates almost always cost the extraction and forget the detection. A scraper that silently returns four hundred rows instead of two thousand does not raise an exception — it returns successfully, with less. Somebody has to notice. Building that noticing is usually larger than building the scraper, and it is the part that gets cut when the deadline moves.",{"type":126,"title":127,"intro":128,"items":129},"list","Signs you have outgrown your own scraper","None of these is fatal on its own. Two or more and the maintenance has become the project.",[130,131,132,133,134],"You have a file called something like fixes.py and nobody remembers what half of it is working around","A redesign on one site has broken your run more than twice this year","You are maintaining a proxy pool, and somebody has had to think about residential versus datacenter IPs","The feed has been wrong in production and the first person to notice was not you","You have started writing per-site exceptions into what was supposed to be a generic parser",{"type":85,"title":136,"intro":137,"headers":138,"rows":142},"Worked example: twelve months of six sites, both ways","The build estimate that loses money is the one that prices the first version and then stops. Six competitor sites, a developer at a loaded cost of €400 a day, and a year of keeping it running. Look at which rows do the damage: the build is 6 days and the year is 20, and the other 14 all arrive after the project was marked done. At one site for one month, build — the maintenance tail never shows up. The tail is the entire decision.",[139,140,141],"","Build it yourself","Buy a web data API",[143,147,151,154,157,161],[144,145,146],"Initial extraction, six sites","6 days — €2,400","1 day of integration — €400",[148,149,150],"Blocking, retries, proxies, scheduling","8 days — €3,200","included",[152,153,150],"Breakages, assuming each site changes twice a year","12 × half a day — €2,400",[155,156,150],"Someone on call to notice a breakage","never in the estimate, and the real cost",[158,159,160],"Twelve-month labour","20 days — €8,000","1 day — €400",[162,163,164],"Infrastructure and fetches","proxies and servers on top of the above","the usage line you compare against",{"type":126,"title":166,"intro":167,"items":168},"What usually goes wrong","Build-versus-buy goes wrong in the estimate far more often than in the engineering.",[169,170,171,172,173],"Pricing the build and forgetting the year. The first working version is 6 of those 20 days; the remaining 14 arrive quietly, one afternoon at a time.","Assuming the retailer has an API. Most official retailer APIs exist to show you your own listings, which is the opposite of what price monitoring needs — a deliberate design decision rather than an oversight.","Counting developer days and ignoring the on-call. The expensive part of a breakage is the eleven days before anyone noticed, not the half day it took to fix.","Comparing a build quote against a list price with no volume attached. The vendor line is pages times frequency; without that multiplication you are comparing a number to a different kind of number.","Treating it as a technical decision. The question is how many sites, for how long, and who is on call when one changes — none of which are about whether you can write the parser.",{"type":80,"title":175,"paragraphs":176},"What a web data API is actually doing for you",[177,178,179],"Three things, and it is worth being precise because the pricing follows directly from them.","It fetches the page, which is the part that involves proxies, headers, retries and occasionally a headless browser. It extracts the fields, which is the part that involves knowing where a price lives on a page that has never been seen before. And it keeps doing both when the site changes, which is the part you are really paying for.","Everything else — scheduling, exports, dashboards — is convenience. If a vendor is strong on the dashboard and vague about what happens when a site redesigns, you are being sold the wrong half.",[181,182,183,184],"Official retailer APIs mostly expose your own listings, not your competitors'. That is by design, not an oversight.","The decision is sites multiplied by duration, not technical difficulty.","Build estimates price the extraction and forget the detection; detection is the expensive half.","What you pay a web data API for is the third thing: still working after the site changes.",{"text":186,"label":187,"url":188},"Next: the first authenticated call, and reading the response without guessing.","Lesson 2: authentication and your first call","/learn/product-data-api/authentication-and-your-first-call",0,[191,192,197,204,210,215],{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":193,"navTitle":194,"title":195,"summary":196,"time":72,"needsAccount":73},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":198,"navTitle":199,"title":200,"summary":201,"time":202,"needsAccount":203},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.","12 min",true,{"slug":205,"navTitle":206,"title":207,"summary":208,"time":209,"needsAccount":73},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.","11 min",{"slug":211,"navTitle":212,"title":213,"summary":214,"time":209,"needsAccount":73},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"slug":216,"navTitle":217,"title":218,"summary":219,"time":202,"needsAccount":203},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",null,{"slug":193,"navTitle":194,"title":195,"summary":196,"time":72,"needsAccount":73},1791047867132]