$17.50 Claude Opus 4.8 per request: $17.50 · per 1M tokens: $11.67
| Model | In $/M | Uit $/M | Kosten | Tokens/$ |
|---|---|---|---|---|
| Claude Opus 4.8 | $5.00 | $25.00 | $17.50 | 40,000 |
| Claude Fable 5 | $10.00 | $50.00 | $35.00 | 20,000 |
Gegevens per 2026-09-27T12:03:11.094Z
(Documentatie in het Engels)
What it does
The LLM Cost Calculator turns per-million-token pricing into real dollars. Two views:
- Per request — input tokens, output tokens, and request count, costed for every model you pick — side by side, cheapest first.
- Monthly workload — “what would 2B tokens a month actually cost me?” Enter tokens/month at billion scale (2B, 500M, plain digits), pick a workload preset for a realistic cache profile, and read your out-of-pocket monthly spend with the cache economics priced lane by lane.
Every rate — input, output, cache read, cache write — flows from one dated pricing snapshot of 85+ models, refreshed on every deploy. Batch API users get the standard 50% discount in one click. Models missing from the pricing table (open weights, unreleased tiers) fall back to a custom-rates mode where you type your own $/M figures.
How to use it
Per-request view
- Set Input tokens / request, Output tokens / request, and Requests with the steppers.
- Click model chips to compare — hover a chip for its live $/M rates; unpriced models are disabled (use custom rates instead).
- Toggle Batch for the 50% discount, or Custom rates for your own per-million figures.
- Read the headline cost with per-request and per-1M-token views; Copy cost or Copy share link.
Monthly workload view
- Switch the top control to Monthly workload.
- Type token volumes per month —
2B,1.5b,500M,10m,2K, or plain digits like1_000_000. - Pick a workload preset — Agent/coding, Chat, RAG, or Batch summarize — each sets a realistic cache-hit rate and cacheable share (tune the two percentages yourself and it flips to Custom).
- Read the monthly total, the blended $/M, and the four-lane table (uncached input / cache reads / cache writes / output) for every selected model, ranked at YOUR hit rate — the model that wins uncached often loses once caching is priced.
Examples
A single large request (per-request view)
1,000,000 input + 500,000 output tokens, 1 request, at $10/M in and $50/M out (Claude Fable 5 list rates):
(1,000,000 / 1M) × $10 + (500,000 / 1M) × $50 = $10 + $25 = $35.00
Batch toggle on: $35.00 × 0.5 = $17.50.
2B tokens/month with agent caching (monthly view)
2B input + 500M output tokens/month, agent/coding preset (75% hit, 95% cacheable), same $10/$50 model with $1 cache reads and $12.5 cache writes per M:
uncached input: 2B × 25% miss /1M × $10 = $ 5,000
cache reads: 2B × 75% hit /1M × $1 = $ 1,500
cache writes: 2B × 25% × 95% /1M × $12.5 = $ 5,625
output: 500M /1M × $50 = $25,000
total = $37,125/month
That is a blended $14.85 per million tokens — versus $45,000/month with caching ignored. Same workload without any caching (hit 0%): $45,000; with a perfect hit rate: $27,000.
Good to know
- Cache writes are priced as write events, not reads. The default derives this month’s write volume from cacheable misses (each cacheable miss is written once, then served by cheaper reads). If you know your real write volume, the lib accepts an explicit override.
- Models without cache pricing degrade, not break — the estimate falls back to no-cache math and the row is flagged (⚠).
- Pricing is a snapshot, refreshed on every deploy — check the “Data as of” line. The site build refreshes it automatically.
- Unpriced models show “—” and cannot be selected; toggle Custom rates and enter your own $/M.
- Private: everything runs 100% client-side — no token counts, workloads, or rates leave your browser.
- CLI twin: the same math ships in the
cosmodevCLI (cosmodev llm-cost-calculator), held to the same test vectors — agents can price workloads locally. - Related tools: Model Picker, Context Window Planner, Cache Savings Calculator.