Pull product data over an API· Lesson 2 of 6

Authentication, keys, and your first real request

Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.

  • 10 min read
  • No account needed

Almost every data API authenticates one of two ways, and the difference matters more than it looks.

The common one is a bearer token in a header. You send Authorization: Bearer followed by the key, the server checks it, and the key never appears in the URL. The other is a key in the query string, which exists because it is trivially easy to test in a browser address bar and is therefore popular in quickstarts.

Prefer the header every time it is offered. A key in a query string ends up in server access logs, in browser history, in the Referer header of any outbound link, in your error-tracking tool's breadcrumb trail, and in the screenshot somebody pastes into a ticket. A key in a header ends up in none of those by default.

The shape of a first request

Nothing vendor-specific here. Replace the base URL and the path with whatever your provider's reference gives you.

bash
curl -s \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  "$API_BASE/scraper/list"

Read the key from the environment, never from the command line directly — your shell history is a plaintext file and it is backed up.

Where the key should live

In rough order of how much trouble each one saves you later.

  • In a secret manager your deployment reads at boot, if you have one
  • In an environment variable set by your orchestrator, if you do not
  • In a gitignored .env file for local development only, with a .env.example that documents the names but holds no values
  • Never in the repository, including in a test fixture, including in a commented-out line, including in a notebook you were going to delete

Reading the first response properly

Two habits here pay for themselves within a week.

First, before you write any parsing code, print the whole response once and read it. Not the field you came for — all of it. Data APIs routinely return metadata alongside the payload: a run identifier, a count, a timestamp, a flag saying the result was truncated. The truncation flag in particular is the kind of thing you find out about by missing it.

Second, check the status code and the body separately. A great many APIs return 200 with an error object inside, because the HTTP request succeeded even though the thing you asked for did not. If your client only branches on the status code, those failures become empty rows rather than exceptions.

Status codes worth branching on

Different responses, genuinely different handling. Treating them all as "error" is why retries get expensive.

CodeWhat it usually meansWhat to do
200 with an error bodyThe request was valid, the operation was notRead the body. Do not retry — it will fail identically.
400Your payload is malformed or a field is unknownFix the request. Retrying is pointless.
401 / 403Key missing, wrong, revoked, or out of scopeStop and alert. A retry loop on a dead key is just noise.
404Wrong path, or an object that no longer existsCheck the reference before assuming the route is gone.
429You are over a rate limitBack off, honour Retry-After if present, and reduce concurrency.
5xxTheir problem, possibly transientRetry with exponential backoff and a cap. See lesson five.

Worked example: three responses that are all 200

Branching on the status code alone is the specific mistake this lesson exists to prevent. All three of these came back 200 OK. One of them is a success.

BodyWhat it actually meansWhat a status-only check does with it
{ "run_id": "r_8812", "rows": [ … 412 rows … ] }A successThe right thing, by luck
{ "run_id": "r_8813", "rows": [] }The run completed and found nothing — a site change, a dead URL list, or a genuinely empty resultWrites an empty table over yesterday's good one
{ "error": "quota_exceeded", "retry_after": 3600 }You are out of creditRecords a success, stores no rows, and tells nobody

What usually goes wrong

Authentication and the first call are where habits get set that are painful to change a year later.

  • Branching on the status and not the body. Two of the three rows above are failures wearing a 200.
  • Putting the key in the query string. It lands in server logs, in browser history, in referrer headers and in every screenshot of the address bar — none of which rotate when you do.
  • One key for everything. When it leaks, and eventually one does, revoking it takes down production, the staging job and somebody's notebook at the same moment.
  • Not establishing the billable unit on day one. Requests, pages fetched and rows returned are three different numbers, and a quote will use whichever is smallest.
  • Testing with a key that has more scope than production will have. The call works in development and 403s on deploy, which is the worst possible time to find out.
  • Retrying a 401. A second attempt changes nothing about a wrong credential, and some providers count the attempts against you.

One thing to verify on day one

Find out, explicitly, what counts as a billable unit. Is it a request you make, a page the provider fetches, or a row you receive? These are not the same number, and the gap between them is where surprise invoices come from.

A run that fetches two thousand pages and returns twelve hundred rows has been billed for two thousand on most pricing models, because the fetch is the cost. If you are budgeting on rows you will be wrong by whatever your failure rate is — and your failure rate is a property of the sites you chose, not of the vendor.

Next: the step most people skip, and the reason their results come back missing half the fields.

Lesson 3: declaring the fields you want

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.