The first thing that surprises people coming from ordinary REST is that starting a collection does not give you any data. It gives you an identifier. The work happens somewhere else, over minutes or hours, and you come back for the result.
This is not an architectural preference. Fetching a single page takes seconds in the best case, and much longer when the site is slow, when a retry is needed, or when a headless browser has to render it. A synchronous endpoint over ten thousand URLs would hold a connection open for hours, and the first network blip would lose the entire result. So the API hands you a job.
The consequence for your code is that there are now three distinct operations where you expected one: start, check, collect.
The three-step shape
Every asynchronous data API is some variation of this, whatever the endpoints are called.
- 1
Start the run
Post the configuration, or the identifier of a saved configuration, and get back a job or run id. Store that id somewhere durable immediately. If your process dies here and the id was only in memory, you have paid for a run you cannot collect.
- 2
Poll until it settles
Ask for the run's state on an interval until it reaches a terminal state. Terminal usually means finished, failed or stopped. Anything else means keep waiting.
- 3
Collect the rows
Fetch the results, usually paginated. This is a separate call and often a separate rate limit, which matters because collecting a large run can take longer than you expect.
How to poll without being rude
A fixed one-second poll on a job that takes forty minutes is two and a half thousand pointless requests, and on some providers it is two and a half thousand requests against your rate limit.
Use exponential backoff with a ceiling. Start at a few seconds, double each time, cap at thirty or sixty seconds. For a job you expect to run for an hour you will make a few dozen calls rather than thousands, and you will still notice completion within a minute of it happening.
If the provider offers a webhook, use it and keep the polling as a fallback rather than deleting it. Webhooks get lost — a deploy restarts your process mid-delivery, a firewall rule changes, the retry policy is less generous than you assumed. A slow poll behind a fast webhook costs almost nothing and converts a silent stall into a late notification.
Pagination and the row cap nobody mentions
Results come back in pages. The mechanics are ordinary — a limit, an offset or a cursor, a loop — but there are two traps.
The first is that sample and preview endpoints frequently cap out at a round number like a hundred rows. If you test with one of those and then build your pagination against it, you will conclude the loop works and that runs are smaller than they are. Check whether the endpoint you are looping over is the full-results endpoint or the preview one.
The second is that the count you use to decide when to stop should come from the run's own metadata, not from counting what you have received. If a page comes back short because of a transient error and your loop uses received-count as its terminator, you will stop early and treat a partial collection as a complete one.
Partial results are the normal case
A run over several thousand URLs will not return several thousand rows. Some pages will be gone, some will be a redirect to a category listing, some will be behind a bot wall on the day you ran it. A ninety per cent return is a good run.
So completion and completeness are different questions, and your code should ask both. The run finished — fine. Did it return roughly what the last one did? That second question is the one that catches real problems, and lesson five of course one covers how to baseline it.
The thing not to do is treat a short run as a failure and retry the whole thing. You will pay twice to collect the same ninety per cent, and the ten per cent that failed will mostly fail again, because the reason was the page and not the attempt.
Terminal states and what each one means for your data
Names vary. The categories do not.
| State | Meaning | Are the rows usable? |
|---|---|---|
| Finished | The run processed its whole input list | Yes, subject to the usual per-page failures |
| Stopped | Something ended it early — you, a cap, or a deadline | Usually yes, but the input was not fully covered |
| Failed | The run itself could not proceed, often a config error | Treat as none; fix the configuration first |
| Finished with errors | Completed, but a meaningful share of pages did not return | Yes, and the error breakdown is the thing to read |
Worked example: a polling schedule that does not hammer
Exponential backoff with a thirty-second cap, against a run that takes four minutes. Twelve requests instead of the 270 a one-second poll would have made, and the longest you ever sit on a finished run is thirty seconds.
| Poll | Wait before it | Elapsed | Status |
|---|---|---|---|
| 1 | 2s | 0:02 | running |
| 2 | 4s | 0:06 | running |
| 3 | 8s | 0:14 | running |
| 4 | 16s | 0:30 | running |
| 5 | 30s (cap reached) | 1:00 | running |
| 6 to 11 | 30s each | 4:00 | running |
| 12 | 30s | 4:30 | succeeded |
What usually goes wrong
Asynchronous collection fails in ways that look like success, which is why the counts matter more than the state.
- Polling on a fixed one-second interval. That is 270 requests where twelve would do, and on a provider with a global rate limit the poller can starve the collection it is waiting for.
- Not storing the run id before the first poll. The process restarts, the id is gone, and the run carries on — and carries on billing — with nobody collecting it.
- No wall-clock deadline. A stalled run polls forever, and forever is a line on an invoice.
- Treating finished as complete. A run can terminate successfully having fetched 900 of 1,200 pages. Ask for the state and the counts, and compare the counts against the last good run.
- Ignoring the row cap. Many endpoints return at most a few hundred rows per page, so a run that produced 4,000 rows hands back 100 and a cursor, and a naive client records 100.
- Throwing away partial results on failure. Eighty per cent of a run is usually worth keeping, provided the gap is recorded rather than quietly absorbed into an average.
Next: retrying properly, including how to avoid paying twice for the same page.
Lesson 5: errors, retries and double billing