| Response headers |
|---|
| x-ratelimit-limit-requests | Your requests-per-minute cap. | x-ratelimit-limit-requests: 60 |
| x-ratelimit-remaining-requests | Requests left in the current window. | x-ratelimit-remaining-requests: 42 |
| x-ratelimit-limit-tokens | Your tokens-per-minute cap (prompt + completion). | x-ratelimit-limit-tokens: 60000 |
| x-ratelimit-remaining-tokens | Tokens left before the next window resets. | x-ratelimit-remaining-tokens: 12340 |
| retry-after | Seconds to wait before retrying (429 responses). | retry-after: 12 |
| Status codes |
|---|
| 429 | Rate limited — back off and retry; safe by definition. | too_many_requests |
| 529 | Provider overloaded — back off; usually transient. | overloaded_error |
| 500 / 503 | Server error — retry with backoff; watch for a pattern. | api_error |
| 400 | Your request is wrong — retrying cannot help. Fix the payload. | invalid_request_error |
| 401 / 403 | Auth failure — check the key; never retry in a loop. | authentication_error |
| Retry strategy |
|---|
| Exponential backoff | Double the delay each attempt: 1s, 2s, 4s, 8s… | delay = base * 2**attempt |
| Jitter | Add randomness so concurrent clients don't retry in lockstep. | delay += random(0, 1s) |
| Respect retry-after | The header beats your computed delay — use the larger value. | delay = max(delay, retry_after) |
| Max attempts | Cap retries (3–5) and surface the error after. | attempts <= 5 |
| Idempotency | Only blind-retry safe operations; regenerate or check otherwise. | retry only GET / validation calls |
| Playing nice |
|---|
| Safety factor | Target ~80% of limits so retries and bursts have headroom. | effective_rpm = rpm * 0.8 |
| Even spacing | Spread requests across the window instead of bursting. | interval = 60s / batch_size |
| Client-side queue | Throttle before the API does — your users see latency, not errors. | p-queue with intervalCap |
| Token budgeting | Large prompts eat TPM faster than RPM — count before you send. | prompt_tokens < remaining_tpm |