Pull product data over an API· Lesson 1 of 6

When an API beats writing your own scraper

Three ways to get product data, the honest cost of each, and the specific question that decides between them.

  • 10 min read
  • No account needed

If you search for "Walmart API documentation" or "Target API key" you will find thousands of other people who searched for the same thing. You will also find that the answer is usually some version of "apply to be a partner", and that the partner API, when you get it, covers your own listings rather than everybody else's. This is not an oversight. A retailer publishes an API so that suppliers can manage their own catalogue; publishing one that lets a competitor read every price would be a strange thing to do on purpose.

So the real choice is not between an official API and something else. It is between three options, and most teams pick one before they have articulated what they are actually optimising for.

The three ways to get product data

Coverage means which sites you can reach. Maintenance means who gets paged when one of them changes.

RouteCoverageMaintenanceBest when
Official retailer APIOne retailer, usually only your own listingsTheirs, and it is versionedYou are a seller on that platform managing your own catalogue
Web data APIAny public page, if the vendor can reach itTheirs, but you still own the schemaYou need several sites and want one integration
Your own scraperWhatever you build, exactlyYours, foreverOne or two sites, stable, and you want full control
A bought datasetBroad but fixed, and already staleNobody's — it is a fileYou need a one-off snapshot for analysis, not a feed

The question that actually decides it

Not "can we build this". You can. A competent developer can read a price off a page in an afternoon, and the first version will work.

The question is: how many distinct sites, and how long does this have to keep working? Those two numbers together are the whole decision. One site for a quarter is a script. Forty sites indefinitely is an operations problem that happens to involve HTTP, and the cost is not the code — it is the on-call rotation, the day a redesign breaks nine of them at once, and the proxy bill you did not budget for.

The honest crossover, from running both: somewhere around three to five sites, or the moment the feed becomes something a person makes a pricing decision from. Below that, write it. Above that, the integration you maintain should be one API rather than forty parsers.

Signs you have outgrown your own scraper

None of these is fatal on its own. Two or more and the maintenance has become the project.

  • You have a file called something like fixes.py and nobody remembers what half of it is working around
  • A redesign on one site has broken your run more than twice this year
  • You are maintaining a proxy pool, and somebody has had to think about residential versus datacenter IPs
  • The feed has been wrong in production and the first person to notice was not you
  • You have started writing per-site exceptions into what was supposed to be a generic parser

Worked example: twelve months of six sites, both ways

The build estimate that loses money is the one that prices the first version and then stops. Six competitor sites, a developer at a loaded cost of €400 a day, and a year of keeping it running. Look at which rows do the damage: the build is 6 days and the year is 20, and the other 14 all arrive after the project was marked done. At one site for one month, build — the maintenance tail never shows up. The tail is the entire decision.

Build it yourselfBuy a web data API
Initial extraction, six sites6 days — €2,4001 day of integration — €400
Blocking, retries, proxies, scheduling8 days — €3,200included
Breakages, assuming each site changes twice a year12 × half a day — €2,400included
Someone on call to notice a breakagenever in the estimate, and the real costincluded
Twelve-month labour20 days — €8,0001 day — €400
Infrastructure and fetchesproxies and servers on top of the abovethe usage line you compare against

What usually goes wrong

Build-versus-buy goes wrong in the estimate far more often than in the engineering.

  • Pricing the build and forgetting the year. The first working version is 6 of those 20 days; the remaining 14 arrive quietly, one afternoon at a time.
  • Assuming the retailer has an API. Most official retailer APIs exist to show you your own listings, which is the opposite of what price monitoring needs — a deliberate design decision rather than an oversight.
  • Counting developer days and ignoring the on-call. The expensive part of a breakage is the eleven days before anyone noticed, not the half day it took to fix.
  • Comparing a build quote against a list price with no volume attached. The vendor line is pages times frequency; without that multiplication you are comparing a number to a different kind of number.
  • Treating it as a technical decision. The question is how many sites, for how long, and who is on call when one changes — none of which are about whether you can write the parser.

What a web data API is actually doing for you

Three things, and it is worth being precise because the pricing follows directly from them.

It fetches the page, which is the part that involves proxies, headers, retries and occasionally a headless browser. It extracts the fields, which is the part that involves knowing where a price lives on a page that has never been seen before. And it keeps doing both when the site changes, which is the part you are really paying for.

Everything else — scheduling, exports, dashboards — is convenience. If a vendor is strong on the dashboard and vague about what happens when a site redesigns, you are being sold the wrong half.

Next: the first authenticated call, and reading the response without guessing.

Lesson 2: authentication and your first call

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.