Skip to content

Cache Savings Calculator — Ruby source

See what prompt caching saves — uncached vs cached cost over N requests, with the write-premium break-even point.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# cache_savings — uncached vs prompt-cached LLM cost comparison.
#
# Language: Ruby (3.2, standard library only)
# Source:   CosmoDev polyglot showcase port of the Cache Savings Calculator
#           tool, ported from src/lib/cacheSavings.ts (the canonical
#           TypeScript implementation).
# Tool:     https://dev.cosmolabs.org/tools/cache-savings-calculator
# License:  display source — part of CosmoDev's polyglot tool pages.
#
# Design goals:
#   - Pure + deterministic; never raises.
#   - Functionally equivalent to the TS reference: same inputs -> same outputs.
#   - Self-contained: core classes only (nil mirrors the TS `null`).
#
# The TS original takes a full AiModel record but reads only its four pricing
# rates, so this port narrows the parameter to exactly those fields. Any nil
# rate makes every output nil — the caller renders an explanatory empty state
# instead of partial math. All rates are per-1M-token USD, mirroring the cost
# conventions of llmCost.ts.
#
# Numeric mapping: TS `number` is a double, so tokens and hits stay Float
# (fractional hits clamp up to 1.0 exactly like `Math.max(1, hits)`);
# breakEvenHits is a whole hit count, so it lands in Integer.

module CacheSavings
  # The four per-1M-token USD pricing rates cache_math reads from the TS
  # AiModel record.
  ModelRates = Struct.new(
    :input_per_m,       # Uncached prompt (input) rate, USD per 1M tokens.
    :output_per_m,      # Completion (output) rate, USD per 1M tokens.
    :cache_read_per_m,  # Cached prompt read rate, USD per 1M tokens.
    :cache_write_per_m, # Cache write premium rate, USD per 1M tokens.
    keyword_init: true
  )

  # Request shape (TS CacheInput).
  CacheInput = Struct.new(
    :prompt_tokens, # Prompt (input) tokens per request.
    :output_tokens, # Completion (output) tokens per request.
    :hits,          # Requests reusing the cached prompt; < 1 counts as 1.
    keyword_init: true
  )

  # Result shape. Any missing rate nils every field.
  CacheMath = Struct.new(
    :uncached,        # hits × (prompt·in$/M + output·out$/M) / 1e6.
    :cached,          # (prompt·write$/M + hits × (prompt·read$/M + output·out$/M)) / 1e6 —
                      # one cache write, `hits` cache reads, output billed every request.
    :savings,         # uncached − cached (negative when caching costs more).
    :savings_pct,     # savings / uncached × 100; 0 when uncached is 0.
    :break_even_hits, # ceil(write$/M / read$/M) when read$/M > 0 — hits needed for
                      # cumulative READ spend to equal ONE write premium; nil otherwise.
    keyword_init: true
  )

  module_function

  # Compare uncached vs prompt-cached cost for one model. Never raises; a nil
  # rate yields the all-nil result so callers can render an explanatory empty
  # state instead of partial math.
  def cache_math(model, input)
    if model.input_per_m.nil? || model.output_per_m.nil? ||
       model.cache_read_per_m.nil? || model.cache_write_per_m.nil?
      return CacheMath.new(uncached: nil, cached: nil, savings: nil,
                           savings_pct: nil, break_even_hits: nil)
    end

    hits = [1.0, input.hits].max
    in_t = input.prompt_tokens.to_f
    out_t = input.output_tokens.to_f
    ipm = model.input_per_m
    opm = model.output_per_m
    cr = model.cache_read_per_m
    cw = model.cache_write_per_m

    # One cache write, `hits` cache reads; output tokens are billed on every request.
    uncached = hits * (in_t * ipm + out_t * opm) / 1_000_000.0
    cached = (in_t * cw + hits * (in_t * cr + out_t * opm)) / 1_000_000.0
    savings = uncached - cached
    savings_pct = uncached.zero? ? 0.0 : savings / uncached * 100.0
    break_even_hits = cr > 0 ? (cw / cr).ceil : nil

    CacheMath.new(uncached: uncached, cached: cached, savings: savings,
                  savings_pct: savings_pct, break_even_hits: break_even_hits)
  end
end

# ----------------------------------------------------------------------
# Self-test — the reference vectors shared with cacheSavings.test.ts (the
# lock-step contract every port mirrors). Run `ruby ruby.rb`.
# ----------------------------------------------------------------------
if $PROGRAM_NAME == __FILE__
  # Fixture model F: input_per_m 10, output_per_m 50, cache_read_per_m 1,
  # cache_write_per_m 12.5. Null one rate to test the unpriced path.
  base = -> { CacheSavings::ModelRates.new(input_per_m: 10.0, output_per_m: 50.0,
                                           cache_read_per_m: 1.0, cache_write_per_m: 12.5) }
  input = ->(p, o, h) { CacheSavings::CacheInput.new(prompt_tokens: p, output_tokens: o, hits: h) }
  close = ->(a, b) { (a - b).abs < 1e-9 }

  # Spec vector: 10k in / 1k out / 5 hits -> uncached 0.75, cached 0.425,
  # savings 0.325, 43.333...% saved, break-even 13 hits.
  r = CacheSavings.cache_math(base.call, input.call(10_000.0, 1_000.0, 5.0))
  raise 'uncached' unless close.call(r.uncached, 0.75)
  raise 'cached' unless close.call(r.cached, 0.425)
  raise 'savings' unless close.call(r.savings, 0.325)
  raise 'savings_pct' unless close.call(r.savings_pct, 43.3333333333)
  raise 'break_even' unless r.break_even_hits == 13 # ceil(12.5 / 1)

  # At 1 hit caching LOSES 0.035 — an honest negative saving.
  r = CacheSavings.cache_math(base.call, input.call(10_000.0, 1_000.0, 1.0))
  raise 'one-hit uncached' unless close.call(r.uncached, 0.15)
  raise 'one-hit cached' unless close.call(r.cached, 0.185)
  raise 'one-hit saving' unless close.call(r.savings, -0.035)
  raise 'one-hit pct' unless close.call(r.savings_pct, -23.3333333333)
  raise 'one-hit break_even' unless r.break_even_hits == 13

  # Each missing rate in turn nils every field.
  [
    base.call.then { |m| m.input_per_m = nil; m },
    base.call.then { |m| m.output_per_m = nil; m },
    base.call.then { |m| m.cache_read_per_m = nil; m },
    base.call.then { |m| m.cache_write_per_m = nil; m }
  ].each do |m|
    x = CacheSavings.cache_math(m, input.call(10_000.0, 1_000.0, 5.0))
    raise 'missing rate must nil everything' unless
      x.uncached.nil? && x.cached.nil? && x.savings.nil? &&
      x.savings_pct.nil? && x.break_even_hits.nil?
  end

  # hits < 1 counts as 1.
  raise 'hits clamp' unless CacheSavings.cache_math(base.call, input.call(10_000.0, 1_000.0, 0.0)) ==
                            CacheSavings.cache_math(base.call, input.call(10_000.0, 1_000.0, 1.0))

  # Zero tokens -> zero costs with 0%, no division error.
  r = CacheSavings.cache_math(base.call, input.call(0.0, 0.0, 5.0))
  raise 'zero tokens' unless r.uncached.zero? && r.cached.zero? &&
                             r.savings.zero? && r.savings_pct.zero?
  raise 'zero-tokens break_even' unless r.break_even_hits == 13

  # cache_read 0 -> break-even nil but costs kept (10k×$12.5 + 5×(0 + 1k×$50)).
  r = CacheSavings.cache_math(base.call.then { |m| m.cache_read_per_m = 0.0; m },
                              input.call(10_000.0, 1_000.0, 5.0))
  raise 'zero read must nil break_even' unless r.break_even_hits.nil? # premium never repaid
  raise 'zero-read uncached' unless close.call(r.uncached, 0.75)
  raise 'zero-read cached' unless close.call(r.cached, 0.375)

  # ceil(4/2) stays 2 — no rounding up at the exact integer boundary.
  r = CacheSavings.cache_math(base.call.then { |m| m.cache_read_per_m = 2.0; m.cache_write_per_m = 4.0; m },
                              input.call(1_000.0, 0.0, 3.0))
  raise 'integer boundary' unless r.break_even_hits == 2

  puts 'ok'
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →