Pull product data over an API· Lesson 5 of 6

Errors, retries, and not paying twice

Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.

  • 11 min read
  • No account needed

Retry logic is where a tidy integration becomes an expensive one. The default instinct — wrap the call in a loop, try three times, move on — is wrong in two directions at once. It retries things that will never succeed, and it retries things that already succeeded.

The second of those is the one that costs money. A request can time out on your side after the server has already done the work. From your perspective nothing happened; from the provider's, a run started and a bill was incurred. Retry and you have two.

Retry or stop

The useful split is not by status code but by whether trying again could plausibly change the outcome.

FailureRetry?Why
Connection reset, DNS blip, timeoutYes, with backoffGenuinely transient. The next attempt is a different network moment.
500, 502, 503, 504Yes, with backoff and a capTheir side, usually brief. Cap it so a long outage does not become an infinite loop.
429Yes, but slowerHonour Retry-After. Also reduce concurrency — retrying at the same rate just re-earns the 429.
400, 422NoThe payload is wrong. It will be wrong on the next attempt too.
401, 403NoThe key is missing, revoked or unscoped. Alert a human instead.
404NoCheck the path against the reference. Retrying a typo is just a slower typo.
200 with an error bodyNoDeterministic. Read the body and handle it as the failure it is.

Backoff needs jitter

Exponential backoff alone has a well-known failure mode. If fifty of your workers all fail at the same moment — which is exactly what happens during a brief provider outage — they all back off by the same amount and all retry at the same instant. You have rebuilt the thundering herd with extra steps, and your retry storm is what keeps the recovery from happening.

Add randomness. Wait somewhere between zero and the backoff interval rather than exactly the interval. It is a one-line change and it is the difference between a recovery and a second outage.

Backoff with jitter, in the smallest honest form

Deliberately boring. The important parts are the cap, the jitter, and that it only loops on retryable failures.

python
import random, time

def with_retries(call, attempts=5, base=2.0, cap=60.0):
    for n in range(attempts):
        try:
            return call()
        except Retryable:
            if n == attempts - 1:
                raise
            delay = min(cap, base * (2 ** n))
            time.sleep(random.uniform(0, delay))

Retryable is your own exception, raised only for the rows in the table above that say yes. Everything else should propagate immediately.

Idempotency keys

An idempotency key is a value you generate and attach to a request that starts work. If the provider supports it, a second request carrying the same key does not start a second run — it returns the result of the first one.

This is the clean answer to the timed-out-but-succeeded problem. Generate the key before the first attempt, reuse it on every retry of that same logical operation, and a retry becomes safe rather than expensive. Generating a fresh key per attempt defeats the entire mechanism, which is a surprisingly common mistake because the generation often lives inside the retry loop by accident.

Two practical notes. Support is usually opt-in, so it does nothing unless you send the header. And it is often supported on more methods than you would guess, including deletes, which matters because a timed-out delete is just as ambiguous as a timed-out create.

Log enough to diagnose it later

What you want in the log line at the moment of failure, because none of it is recoverable afterwards.

  • The run or request id the provider gave you — without it, support cannot help and neither can you
  • The status code and the first part of the response body, not just the exception message
  • Which attempt number this was, so a one-off is distinguishable from a systematic failure
  • The timestamp in UTC, because the provider's logs are in UTC and correlating across zones at 3am is how mistakes happen
  • A durable note of whether work may have started, so the next operator knows whether a retry is safe

Worked example: how one timeout becomes three invoices

A start-run request times out at the client after thirty seconds. The server received it and is working on it. Without an idempotency key every retry is a new run, and all of them finish. With one key generated before the first attempt, attempts two and three return run A's id and the bill is one run — but the key has to exist before attempt one. A key generated inside the retry loop is a new key each time and does nothing whatsoever.

AttemptWhat the client seesWhat the server doesRuns in flight
1timeout after 30saccepted, starts run A1
2, a retrytimeout after 30saccepted, starts run B2
3, a retry200, run id Caccepted, starts run C3
Outcomeone run idthree complete runsthree times the pages, a third of them collected

What usually goes wrong

Retry logic is written on a good day and exercised on a bad one, which is why these are so common.

  • Retrying a 400. Nothing about a second attempt changes a malformed request, and some providers charge for the attempt anyway.
  • Backoff without jitter. Every client that failed during the outage retries on exactly the same schedule and rebuilds the stampede at the moment the service is trying to recover.
  • Generating the idempotency key inside the retry loop. A fresh key per attempt is identical to no key, and it reads as correct in the code.
  • Assuming the rate limit is per endpoint. It is frequently global, so an aggressive poller on one route can 429 the route that actually matters.
  • Retrying a quota or payment error. That is the class that means stop: the system is working correctly and telling you that you are out of credit.
  • Logging that a call failed without logging what was sent. A month later you have a count of failures and no way to reproduce a single one.

The failure that means stop, not try harder

There is one category where retrying is not just useless but actively the wrong instinct: when the page is refusing you rather than failing.

A bot wall that has decided your traffic is automated does not give you a flaky response that works on the third attempt. It gives you a consistent refusal, and hammering it makes the classification worse. Three retries against a wall is three times the cost for the same answer, plus a stronger signal that you are what the wall thinks you are.

This is the main subject of the next course, and the distinction worth carrying into it is between a page that failed and a page that declined. Your retry logic can only help with the first.

Next: getting the feed out of the API and into something your team actually queries.

Lesson 6: putting the feed into your stack

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.