Free (~60 RPM / 60k TPM) (preset — edit freely)
Prompt + completion. Not sure? 1 token ≈ 4 characters of English.
Headroom left for retries and bursts — 80% is a sane default.
What it does
The Rate Limit Planner turns a provider’s rate limits (RPM — requests per minute, TPM — tokens per minute) and your workload into a schedule you can actually run: how many requests per 60-second window, how far apart to space them, which limit binds first, and when the whole batch finishes. A safety factor (default 80%) leaves headroom for retries so one 429 doesn’t cascade. Runs 100% client-side.
How to use it
- Pick a provider tier preset (or Custom) and adjust RPM/TPM to your provider’s documented limits.
- Enter the total requests and the average tokens per request (prompt + completion).
- Tune the safety factor — 80% is a sane default; use 90–100% only when you retry aggressively on 429s.
- Read the plan: batch size per window, spacing, binding limit, total time, and the first ten windows of the schedule. Copy the schedule for your runbook.
Examples
60 RPM / 60k TPM, 200 requests averaging 1,200 tokens, 80% safety:
RPM-bound? min(48, 48000·0.8/1200 = 32) → 32 requests/window (TPM binds)
spacing 60000 / 32 ≈ 1875 ms
total ⌈200/32⌉ = 7 windows → ~6 min 9 s
500 RPM / 300k TPM, 50 requests averaging 500 tokens:
batch min(400, 480) → 400/window (RPM binds) — finishes inside one window
Good to know
- Token-bound plans are common: large prompts eat TPM far faster than RPM. If one request’s average exceeds your effective TPM, no schedule can run it — the planner says so instead of inventing one.
- The schedule assumes even spacing inside each 60s window. If your client only supports fixed bursts, treat “batch size” as the burst and wait for the window to roll over.
- Tier presets are rough public figures for orientation — always confirm against your provider’s current docs.
- Runs 100% client-side. Nothing is sent anywhere.
- Related: Token Estimator, Context Window Planner, LLM Cost Calculator.