Skip to content

Rate Limits & Retries Explained

Cheatsheet of LLM API rate limiting: the response headers to read, the error codes to expect, and the retry strategies that actually work.

Every LLM API throttles by requests, tokens, or both. The providers differ in the numbers but agree on the mechanics: read the headers, back off on 429, and never retry a non-idempotent failure blind.

Reference table · 19 entriesOpen the rate-limit-planner tool →
19 of 19 rows
Response headers
x-ratelimit-limit-requestsYour requests-per-minute cap.x-ratelimit-limit-requests: 60
x-ratelimit-remaining-requestsRequests left in the current window.x-ratelimit-remaining-requests: 42
x-ratelimit-limit-tokensYour tokens-per-minute cap (prompt + completion).x-ratelimit-limit-tokens: 60000
x-ratelimit-remaining-tokensTokens left before the next window resets.x-ratelimit-remaining-tokens: 12340
retry-afterSeconds to wait before retrying (429 responses).retry-after: 12
Status codes
429Rate limited — back off and retry; safe by definition.too_many_requests
529Provider overloaded — back off; usually transient.overloaded_error
500 / 503Server error — retry with backoff; watch for a pattern.api_error
400Your request is wrong — retrying cannot help. Fix the payload.invalid_request_error
401 / 403Auth failure — check the key; never retry in a loop.authentication_error
Retry strategy
Exponential backoffDouble the delay each attempt: 1s, 2s, 4s, 8s…delay = base * 2**attempt
JitterAdd randomness so concurrent clients don't retry in lockstep.delay += random(0, 1s)
Respect retry-afterThe header beats your computed delay — use the larger value.delay = max(delay, retry_after)
Max attemptsCap retries (3–5) and surface the error after.attempts <= 5
IdempotencyOnly blind-retry safe operations; regenerate or check otherwise.retry only GET / validation calls
Playing nice
Safety factorTarget ~80% of limits so retries and bursts have headroom.effective_rpm = rpm * 0.8
Even spacingSpread requests across the window instead of bursting.interval = 60s / batch_size
Client-side queueThrottle before the API does — your users see latency, not errors.p-queue with intervalCap
Token budgetingLarge prompts eat TPM faster than RPM — count before you send.prompt_tokens < remaining_tpm