Skip to content

VRAM Calculator — Ruby source

Estimate the VRAM an LLM needs — weights by quantization plus the KV cache for your context and batch — and see which consumer and datacenter GPUs hold it.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# vram-calculator — Ruby port: estimate the VRAM an LLM needs (weights + KV cache).

# Bytes stored per weight for each quantization (Q4_K_M = 4.85 bits/weight,
# the llama.cpp mix). Sizes are decimal gigabytes, matching how parameter
# counts ("70B") and GPU sizes are quoted.
QUANT_BYTES = { fp32: 4.0, fp16: 2.0, bf16: 2.0, int8: 1.0, int4: 0.5, q4_K_M: 4.85 / 8 }.freeze

# Default attention architecture: a modern GQA-style layout. Override per
# model — e.g. Llama-2-70B uses 80 layers with the same GQA shape.
DEFAULT_ARCH = { layers: 32, kv_heads: 8, head_dim: 128 }.freeze
DEFAULT_KV_BYTES = 2 # bytes per KV element (fp16 K and V tensors)

GB = 1e9

# Estimate the VRAM footprint: weights plus KV cache, both in decimal GB.
#   weights = params_b * bytes_per_param
#   kv      = 2 * layers * context * kv_heads * head_dim * kv_bytes * batch / 1e9
# Raises ArgumentError on params_b <= 0, unknown quantization, negative
# context, or any option below 1 (context 0 allowed — no context, no cache).
def vram(params_b, quant, context, layers: DEFAULT_ARCH[:layers], kv_heads: DEFAULT_ARCH[:kv_heads],
         head_dim: DEFAULT_ARCH[:head_dim], batch: 1, kv_bytes: DEFAULT_KV_BYTES)
  raise ArgumentError, "params_b must be finite > 0 (got #{params_b})" unless params_b.finite? && params_b.positive?

  bytes_per_param = QUANT_BYTES[quant] ||
                    raise(ArgumentError, "unknown quantization #{quant} — expected one of #{QUANT_BYTES.keys.join(', ')}")
  raise ArgumentError, "context must be >= 0 (got #{context})" if context.negative?

  { layers: layers, kv_heads: kv_heads, head_dim: head_dim, batch: batch, kv_bytes: kv_bytes }.each do |name, v|
    raise ArgumentError, "#{name} must be >= 1 (got #{v})" unless v.is_a?(Numeric) && v >= 1
  end

  weights_gb = params_b * 1e9 * bytes_per_param / GB
  kv_cache_gb = 2 * layers * context * kv_heads * head_dim * kv_bytes * batch / GB
  { quant: quant, bytes_per_param: bytes_per_param, weights_gb: weights_gb,
    kv_cache_gb: kv_cache_gb, total_gb: weights_gb + kv_cache_gb }
end

# Common GPU memory tiers, from consumer boards to datacenter cards.
GPU_CARDS = [
  { name: 'RTX 3060 Ti / RTX 4060 / RX 7600', size_gb: 8 },
  { name: 'RTX 3060 12 GB / RTX 4070', size_gb: 12 },
  { name: 'RTX 4060 Ti 16 GB / RTX 5080', size_gb: 16 },
  { name: 'RTX 3090 / RTX 4090', size_gb: 24 },
  { name: 'RTX A6000 / L40S', size_gb: 48 },
  { name: 'A100 80 GB / H100 / H200', size_gb: 80 }
].freeze

# Score every card against a total. `fits` is inclusive: a total exactly
# equal to the card size fits (headroom 0).
def gpu_fits(total_gb, cards = GPU_CARDS)
  raise ArgumentError, "total_gb must be finite >= 0 (got #{total_gb})" unless total_gb.finite? && total_gb >= 0

  cards.map do |c|
    headroom = c[:size_gb] - total_gb
    c.merge(fits: headroom >= 0, headroom_gb: headroom)
  end
end

# Demo — the lib's canonical vectors: a 70B Llama-2-shape build (80 layers)
# and an 8B default-arch build, then the GPU fit table for the 8B total.
r70 = vram(70, :q4_K_M, 4096, layers: 80)
r8 = vram(8, :fp16, 8192)
puts format('70B q4_K_M @4096 (80 layers): %.2f GB weights + %.2f GB KV = %.2f GB total',
            r70[:weights_gb], r70[:kv_cache_gb], r70[:total_gb])
puts format('8B  fp16    @8192:            %.2f GB weights + %.2f GB KV = %.2f GB total',
            r8[:weights_gb], r8[:kv_cache_gb], r8[:total_gb])
puts format('GPU fit for %.2f GB:', r8[:total_gb])
gpu_fits(r8[:total_gb]).each do |f|
  puts format('  %-36s %s (%.2f GB %s)', f[:name], f[:fits] ? 'fits' : 'too small',
              f[:headroom_gb].abs, f[:fits] ? 'headroom' : 'short')
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →