Paste your prompt as labeled sections — system prompt, retrieved docs, conversation history. Each paste is token-estimated client-side, then laid against every selected model’s context window. Pick models above to see fill bars, free headroom, and overflow warnings.
What it does
The Context Window Planner shows how a real prompt fills a model’s context window. Paste the pieces of your request as labeled sections — system prompt, retrieved docs, conversation history — and the tool token-estimates each one, sums them, and lays the total against every selected model’s window as a fill bar.
It answers the two questions that break LLM requests in production: does this prompt even fit? and if it fits, how much room is left for the answer? Set an output reserve and each model reports whether the free headroom covers it, so a prompt that technically fits but starves the response gets flagged before you ship it.
How to use it
- Paste your System prompt, Docs, and History into the stacked sections. Add or remove sections to match your real prompt shape.
- Set the Output reserve — the tokens you want kept free for the model’s reply (a common cause of truncated output is a reserve of zero).
- Toggle the model chips (widest windows first) to compare fills across vendors and tiers.
- Read the fill bars: blue while the input fits, red with the overflow amount when it does not; “reserve short” marks inputs that fit but leave less headroom than your reserve.
- Copy summary or Copy share link to send the exact scenario.
Examples
A two-section prompt against a 1M-token window
Each section below is one 1,600-character line of prose (≈400 tokens at 4 chars/token):
System: 1,600 chars → 400 tokens
Docs: 1,600 chars → 400 tokens
beta-pro (1,000,000 ctx): 800 / 1,000,000 · free 999,200 · fits yes · reserve ok
The same prompt that fits, but starves the reserve
input 800 · output reserve 1,000,000
beta-pro: fits yes · reserve SHORT ← raw fit passes, no room for the answer
An overflow, flagged with the amount
A single 4.4M-character paste (≈1.1M tokens) against a 1M window:
beta-pro: 1,100,000 / 1,000,000 · overflow −100,000 · fits NO
Good to know
- Fitting the window is not the same as fitting the answer.
fitsmeans the raw input is within the context window; the reserve check means there is still room for the response. Both must pass for a healthy request. - Token counts are heuristic estimates (±15%), classified per line as prose, code, JSON, or CJK — not a real BPE tokenizer run. For the estimate breakdown alone, see the Token Estimator.
- Model specs are a snapshot. Window sizes and output caps flow from one dated model snapshot — check the “Data as of” line for freshness.
- Private: everything runs 100% client-side — no prompt text ever leaves your browser.
- Related tools: Token Estimator, Model Picker, LLM Cost Calculator.