صرفهجویی $0.33 (43.33%) $0.75 بدون کش → $0.43 با کش · 5 درخواست سربهسر: 13 اصابت کش — خواندنها باید یک حق نوشتن اضافه را جبران کنند
بدون کش = hits × (prompt×in$/M + output×out$/M) / 1M
= 5 × (10,000×$10 + 1,000×$50) / 1M = $0.75
با کش = (prompt×write$/M + hits × (prompt×read$/M + output×out$/M)) / 1M
= (10,000×$12.5 + 5 × (10,000×$1 + 1,000×$50)) / 1M = $0.43
break-even hits = ceil(write$/M / read$/M) = ceil($12.5 / $1) = 13داده تا 2026-09-27T12:03:11.094Z
(مستندات به انگلیسی)
What it does
The Cache Savings Calculator shows what prompt caching actually saves on a repeated workload. Enter your prompt and output token counts plus how many requests reuse the cached prompt, pick a model, and it compares the uncached bill against the cached one — including the honest case where caching costs more.
The key nuance it captures is the write premium: the first request pays extra to store the prompt, and every later hit pays a much cheaper cache-read rate instead of full input price. The break-even figure tells you how many hits it takes for the cheap reads to repay that one write.
How to use it
- Set Prompt tokens / request, Output tokens / request, and Cache hits with the steppers.
- Pick a Model — cache-priced models sort first; models without published cache rates are marked “no cache rates” and show the explanatory empty state.
- Read the headline: green when caching saves money, red when it costs more (too few hits to amortize the write).
- Check The math block — both formulas rendered with your model’s live rates and your numbers substituted in.
- Copy savings or Copy share link to send the exact scenario.
Examples
A repeated 10k-token prompt (5 hits)
Input: 10,000 prompt tokens, 1,000 output tokens, 5 cache hits, at $10/M input, $50/M output, $1/M cache read, $12.5/M cache write:
uncached = 5 × (10,000×$10 + 1,000×$50) / 1M = $0.75
cached = (10,000×$12.5 + 5 × (10,000×$1 + 1,000×$50)) / 1M = $0.425
savings = $0.325 (43.33%) · break-even = ceil(12.5 / 1) = 13 hits
A single request — caching loses
Same shape, 1 cache hit:
uncached = $0.15 · cached = $0.185 · savings = −$0.035
One hit never amortizes the write premium; the tool shows the negative result in an error tone rather than hiding it.
Good to know
- Break-even counts reads, not requests.
ceil(write$/M ÷ read$/M)is how many cache hits it takes for cumulative read spend to equal one write premium — below it, caching costs more. - Rates are a snapshot. Every figure flows from one dated model snapshot — check the “Data as of” line for freshness; models without published cache rates return no result instead of partial math.
- Private: everything runs 100% client-side — no token counts or rates leave your browser.
- Related tools: LLM Cost Calculator, Token Estimator, Embedding Chunk Planner.