Skip to content

VRAM Calculator

Estimate the VRAM an LLM needs — weights by quantization plus the KV cache for your context and batch — and see which consumer and datacenter GPUs hold it.

AI & Agents
cosmodev ~/tools/vram-calculator-
Params:
14 GB
Weights · fp16 · 2 B/weight
1.07 GB
KV cache · batch 1
15.07 GB
Total VRAM · weights + KV cache
Card classVRAMFree after loadStatus
RTX 3060 Ti / RTX 4060 / RX 76008 GB—Too small
RTX 3060 12 GB / RTX 407012 GB—Too small
RTX 4060 Ti 16 GB / RTX 508016 GB0.93 GBFits
RTX 3090 / RTX 409024 GB8.93 GBFits
RTX A6000 / L40S48 GB32.93 GBFits
A100 80 GB / H100 / H20080 GB64.93 GBFits

weights = params × bytes/weight · KV = 2 × layers × context × kvHeads × headDim × 2 B × batch · decimal GB (10⁹ bytes); activations and runtime overhead not included