Each message: { role, tokens, pinned? }. "pinned": true is never pruned.
Context window minus what you reserve for the reply. history: 1,103 tok
What it does
The Conversation Pruner turns a long chat history and a token budget into an explicit pruning plan: every message marked keep, summarize, or drop. System messages, pinned turns, the opening request, and the current message are never pruned; recent turns are kept newest-first until the budget fills; the squeezed middle folds into one running summary when the compressed form fits, and drops when it doesn’t. Runs 100% client-side.
How to use it
- Paste your conversation as a JSON array — one object per message:
{ "role": "user", "tokens": 120 }. Add"pinned": trueto protect a turn. - Set the history budget: your context window minus what you reserve for the reply.
- Read the plan: the narrative line, any warnings, and the per-message verdict list. Copy the plan into your runbook or client code.
Examples
A 1,090-token chat against a 1,000-token budget:
#0 system KEEP 220 (system — protected)
#1 user KEEP 100 (opening request — protected)
#2 assistant SUMMARIZE 310 ┐
#3 user SUMMARIZE 130 ┘ folded into a 90-token summary
#4 assistant KEEP 260 (newest turn that fits)
#5 user KEEP 88 (current request — protected)
Projected: 880 tokens — fits. Against a 700-token budget the same fold group no longer fits compressed, so turns #2–#3 drop instead and the plan still lands at 670.
Good to know
- The summary cost model is explicit: 60 fixed tokens + 10% of the folded content — a 400-token middle becomes a ~100-token summary. Swap in your own numbers once you know your model’s summary behavior.
- “Fits” counts only what the plan keeps plus an applied summary; a rejected summary costs nothing.
- Token counts per message are yours to supply (prompt + framing). The Token Estimator can produce them from raw text.
- Runs 100% client-side. Your transcript never leaves the browser.
- Related: Token Estimator, Context Window Planner, Rate Limit Planner.