VRAM Calculator — Ruby source
Estimate the VRAM an LLM needs — weights by quantization plus the KV cache for your context and batch — and see which consumer and datacenter GPUs hold it.
This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.
# vram-calculator — Ruby port: estimate the VRAM an LLM needs (weights + KV cache).
# Bytes stored per weight for each quantization (Q4_K_M = 4.85 bits/weight,
# the llama.cpp mix). Sizes are decimal gigabytes, matching how parameter
# counts ("70B") and GPU sizes are quoted.
QUANT_BYTES = { fp32: 4.0, fp16: 2.0, bf16: 2.0, int8: 1.0, int4: 0.5, q4_K_M: 4.85 / 8 }.freeze
# Default attention architecture: a modern GQA-style layout. Override per
# model — e.g. Llama-2-70B uses 80 layers with the same GQA shape.
DEFAULT_ARCH = { layers: 32, kv_heads: 8, head_dim: 128 }.freeze
DEFAULT_KV_BYTES = 2 # bytes per KV element (fp16 K and V tensors)
GB = 1e9
# Estimate the VRAM footprint: weights plus KV cache, both in decimal GB.
# weights = params_b * bytes_per_param
# kv = 2 * layers * context * kv_heads * head_dim * kv_bytes * batch / 1e9
# Raises ArgumentError on params_b <= 0, unknown quantization, negative
# context, or any option below 1 (context 0 allowed — no context, no cache).
def vram(params_b, quant, context, layers: DEFAULT_ARCH[:layers], kv_heads: DEFAULT_ARCH[:kv_heads],
head_dim: DEFAULT_ARCH[:head_dim], batch: 1, kv_bytes: DEFAULT_KV_BYTES)
raise ArgumentError, "params_b must be finite > 0 (got #{params_b})" unless params_b.finite? && params_b.positive?
bytes_per_param = QUANT_BYTES[quant] ||
raise(ArgumentError, "unknown quantization #{quant} — expected one of #{QUANT_BYTES.keys.join(', ')}")
raise ArgumentError, "context must be >= 0 (got #{context})" if context.negative?
{ layers: layers, kv_heads: kv_heads, head_dim: head_dim, batch: batch, kv_bytes: kv_bytes }.each do |name, v|
raise ArgumentError, "#{name} must be >= 1 (got #{v})" unless v.is_a?(Numeric) && v >= 1
end
weights_gb = params_b * 1e9 * bytes_per_param / GB
kv_cache_gb = 2 * layers * context * kv_heads * head_dim * kv_bytes * batch / GB
{ quant: quant, bytes_per_param: bytes_per_param, weights_gb: weights_gb,
kv_cache_gb: kv_cache_gb, total_gb: weights_gb + kv_cache_gb }
end
# Common GPU memory tiers, from consumer boards to datacenter cards.
GPU_CARDS = [
{ name: 'RTX 3060 Ti / RTX 4060 / RX 7600', size_gb: 8 },
{ name: 'RTX 3060 12 GB / RTX 4070', size_gb: 12 },
{ name: 'RTX 4060 Ti 16 GB / RTX 5080', size_gb: 16 },
{ name: 'RTX 3090 / RTX 4090', size_gb: 24 },
{ name: 'RTX A6000 / L40S', size_gb: 48 },
{ name: 'A100 80 GB / H100 / H200', size_gb: 80 }
].freeze
# Score every card against a total. `fits` is inclusive: a total exactly
# equal to the card size fits (headroom 0).
def gpu_fits(total_gb, cards = GPU_CARDS)
raise ArgumentError, "total_gb must be finite >= 0 (got #{total_gb})" unless total_gb.finite? && total_gb >= 0
cards.map do |c|
headroom = c[:size_gb] - total_gb
c.merge(fits: headroom >= 0, headroom_gb: headroom)
end
end
# Demo — the lib's canonical vectors: a 70B Llama-2-shape build (80 layers)
# and an 8B default-arch build, then the GPU fit table for the 8B total.
r70 = vram(70, :q4_K_M, 4096, layers: 80)
r8 = vram(8, :fp16, 8192)
puts format('70B q4_K_M @4096 (80 layers): %.2f GB weights + %.2f GB KV = %.2f GB total',
r70[:weights_gb], r70[:kv_cache_gb], r70[:total_gb])
puts format('8B fp16 @8192: %.2f GB weights + %.2f GB KV = %.2f GB total',
r8[:weights_gb], r8[:kv_cache_gb], r8[:total_gb])
puts format('GPU fit for %.2f GB:', r8[:total_gb])
gpu_fits(r8[:total_gb]).each do |f|
puts format(' %-36s %s (%.2f GB %s)', f[:name], f[:fits] ? 'fits' : 'too small',
f[:headroom_gb].abs, f[:fits] ? 'headroom' : 'short')
end
Also available in 13 other languages
Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →