Two or more real requests from your app — the structure they share is what caches.
What it does
Prompt caching only pays on the parts of a prompt that stay identical across requests. The Cache Breakpoint Planner compares your real prompt sessions block by block, finds the shared prefix (and shared tail), tells you exactly where to place cache breakpoints, and estimates the cost saving from the ~10× cached-read discount providers like Anthropic document. Runs 100% client-side.
How to use it
- Paste two or more real requests from your app as JSON —
{ "id": "session-a", "blocks": ["system…", "docs…", "user…"] }. Split blocks where your code assembles the prompt. - Read the verdict: the shared prefix/suffix block counts, a breakpoint placement after the last shared block, and the estimated saving.
- Restructure until the cacheable share is as high as your product allows: stable content first, per-request content last.
Examples
Two helpdesk requests sharing system + policy blocks:
shared prefix: 2 blocks, ~55 tokens
breakpoint after block 2 — cache once, hit on every request
estimated saving vs no cache: ~74%
Move the volatile user line before the policy block and the prefix drops to one block — the planner shows the saving fall immediately, which is the whole lesson.
Good to know
- Block order is the product decision: whatever varies (user text, timestamps, retrieved chunks) belongs after the last breakpoint.
- The saving estimate assumes cached reads bill at 0.1× input price and every session after the first hits the cache — real savings land slightly below it.
- A shared tail can be folded into the cached segment only if it is also stable; otherwise it re-reads at full price every request.
- Token figures are estimates (prose heuristic, ~15% tolerance).
- Runs 100% client-side. Your prompts never leave the browser.
- Related: System Prompt Builder, Token Estimator, LLM Cost Calculator.