- Published on
Scraping: retry with backoff, honour Retry-After, and know what not to retry
- Authors

- Name
- Nadim Tuhin
- @nadimtuhin
On this page
A scraper that stops on the first failed request is brittle. One that retries everything is worse, because it hammers a server that already told it to slow down.
The rule I use is short:
- Retry the failures that can clear on their own: 429 and 5xx, plus connection errors and timeouts.
- Do not retry the rest. A 404 will still be a 404 in a second.
- If the server says how long to wait, believe it.
The helper
import random, time, requests
RETRY = {429, 500, 502, 503, 504}
def get(url, tries=5, base=0.2, cap=5):
for n in range(1, tries + 1):
r = requests.get(url, timeout=(3, 10))
if r.status_code not in RETRY:
return r, n
ra = r.headers.get("Retry-After")
wait = float(ra) if ra and ra.isdigit() else min(cap, base * 2**n) * random.uniform(0.5, 1.0)
print(f" attempt {n}: {r.status_code}, waiting {wait:.2f}s")
time.sleep(wait)
r.raise_for_status()
Three details matter:
- Timeouts are a tuple.
(3, 10)is 3 seconds to connect and 10 to read. With no timeout,requestswaits forever. - Jitter. Multiplying by a random factor stops many workers retrying in lockstep.
- A cap. Without
cap, the wait doubles into minutes.
What it did
I pointed it at a local server that returns 429 with Retry-After: 1 on the first call, 503 on the second, and 200 on the third. A second URL always returns 404:
attempt 1: 429, waiting 1.00s
attempt 2: 503, waiting 0.73s
flaky -> 200 {'ok': True, 'attempt': 3} after 3 attempts in 1.74 s
missing -> 404 after 1 attempt in 0.0 s
The 429 wait was exactly the 1 second the server asked for. The 503 had no header, so it used the jittered backoff. The 404 returned at once.
Limits of this sketch
- It does not parse a
Retry-Afterthat is an HTTP date, only seconds. - It retries
GETonly. Do not retry aPOSTunless the endpoint is idempotent. - Retries do nothing about your request rate. If you get a 429 often, send fewer requests.