Almost every data API authenticates one of two ways, and the difference matters more than it looks.
The common one is a bearer token in a header. You send Authorization: Bearer followed by the key, the server checks it, and the key never appears in the URL. The other is a key in the query string, which exists because it is trivially easy to test in a browser address bar and is therefore popular in quickstarts.
Prefer the header every time it is offered. A key in a query string ends up in server access logs, in browser history, in the Referer header of any outbound link, in your error-tracking tool's breadcrumb trail, and in the screenshot somebody pastes into a ticket. A key in a header ends up in none of those by default.
The shape of a first request
Nothing vendor-specific here. Replace the base URL and the path with whatever your provider's reference gives you.
curl -s \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
"$API_BASE/scraper/list"Read the key from the environment, never from the command line directly — your shell history is a plaintext file and it is backed up.
Where the key should live
In rough order of how much trouble each one saves you later.
- In a secret manager your deployment reads at boot, if you have one
- In an environment variable set by your orchestrator, if you do not
- In a gitignored .env file for local development only, with a .env.example that documents the names but holds no values
- Never in the repository, including in a test fixture, including in a commented-out line, including in a notebook you were going to delete
Reading the first response properly
Two habits here pay for themselves within a week.
First, before you write any parsing code, print the whole response once and read it. Not the field you came for — all of it. Data APIs routinely return metadata alongside the payload: a run identifier, a count, a timestamp, a flag saying the result was truncated. The truncation flag in particular is the kind of thing you find out about by missing it.
Second, check the status code and the body separately. A great many APIs return 200 with an error object inside, because the HTTP request succeeded even though the thing you asked for did not. If your client only branches on the status code, those failures become empty rows rather than exceptions.
Status codes worth branching on
Different responses, genuinely different handling. Treating them all as "error" is why retries get expensive.
| Code | What it usually means | What to do |
|---|---|---|
| 200 with an error body | The request was valid, the operation was not | Read the body. Do not retry — it will fail identically. |
| 400 | Your payload is malformed or a field is unknown | Fix the request. Retrying is pointless. |
| 401 / 403 | Key missing, wrong, revoked, or out of scope | Stop and alert. A retry loop on a dead key is just noise. |
| 404 | Wrong path, or an object that no longer exists | Check the reference before assuming the route is gone. |
| 429 | You are over a rate limit | Back off, honour Retry-After if present, and reduce concurrency. |
| 5xx | Their problem, possibly transient | Retry with exponential backoff and a cap. See lesson five. |
Worked example: three responses that are all 200
Branching on the status code alone is the specific mistake this lesson exists to prevent. All three of these came back 200 OK. One of them is a success.
| Body | What it actually means | What a status-only check does with it |
|---|---|---|
| { "run_id": "r_8812", "rows": [ … 412 rows … ] } | A success | The right thing, by luck |
| { "run_id": "r_8813", "rows": [] } | The run completed and found nothing — a site change, a dead URL list, or a genuinely empty result | Writes an empty table over yesterday's good one |
| { "error": "quota_exceeded", "retry_after": 3600 } | You are out of credit | Records a success, stores no rows, and tells nobody |
What usually goes wrong
Authentication and the first call are where habits get set that are painful to change a year later.
- Branching on the status and not the body. Two of the three rows above are failures wearing a 200.
- Putting the key in the query string. It lands in server logs, in browser history, in referrer headers and in every screenshot of the address bar — none of which rotate when you do.
- One key for everything. When it leaks, and eventually one does, revoking it takes down production, the staging job and somebody's notebook at the same moment.
- Not establishing the billable unit on day one. Requests, pages fetched and rows returned are three different numbers, and a quote will use whichever is smallest.
- Testing with a key that has more scope than production will have. The call works in development and 403s on deploy, which is the worst possible time to find out.
- Retrying a 401. A second attempt changes nothing about a wrong credential, and some providers count the attempts against you.
One thing to verify on day one
Find out, explicitly, what counts as a billable unit. Is it a request you make, a page the provider fetches, or a row you receive? These are not the same number, and the gap between them is where surprise invoices come from.
A run that fetches two thousand pages and returns twelve hundred rows has been billed for two thousand on most pricing models, because the fetch is the cost. If you are budgeting on rows you will be wrong by whatever your failure rate is — and your failure rate is a property of the sites you chose, not of the vendor.
Next: the step most people skip, and the reason their results come back missing half the fields.
Lesson 3: declaring the fields you want