| کلاس کارت | VRAM | آزاد پس از بارگذاری | وضعیت |
|---|---|---|---|
| RTX 3060 Ti / RTX 4060 / RX 7600 | 8 GB | — | خیلی کوچک است |
| RTX 3060 12 GB / RTX 4070 | 12 GB | — | خیلی کوچک است |
| RTX 4060 Ti 16 GB / RTX 5080 | 16 GB | 0.93 GB | جا میشود |
| RTX 3090 / RTX 4090 | 24 GB | 8.93 GB | جا میشود |
| RTX A6000 / L40S | 48 GB | 32.93 GB | جا میشود |
| A100 80 GB / H100 / H200 | 80 GB | 64.93 GB | جا میشود |
وزنها = params × bytes/weight · KV = 2 × layers × context × kvHeads × headDim × 2 B × batch · GB دهدهی (10⁹ بایت)؛ فعالسازیها و سربار زمان اجرا شامل نمیشود
(مستندات به انگلیسی)
What it does
The VRAM Calculator estimates how much GPU memory a large language model needs before you download or rent anything. Enter the parameter count and its quantization (FP16, INT8, Q4_K_M, …) and it computes the weights footprint; add your context length, batch size, and the model’s attention shape and it adds the KV-cache cost of actually serving that context. It then checks the total against common GPU memory tiers — from an 8 GB consumer card to an 80 GB datacenter board — showing how much stays free after the model loads. Everything runs 100% client-side.
How to use it
- Pick a parameter preset (0.5B … 70B) or type any value into Parameters (billions).
- Choose a Quantization — the bytes-per-weight (4, 2, 1, 0.5, or ~0.61 for Q4_K_M) drives the weights size.
- Set Context length and Batch size — both multiply the KV cache.
- If you know the model’s shape, override Layers, KV heads (after GQA), and Head dimension under Attention architecture; the defaults (32 / 8 / 128) fit a modern GQA model.
- Read the Weights / KV cache / Total tiles, then check the GPU fit table for which cards hold it and how much VRAM stays free.
- Copy total for the headline number, or Copy share link to send the exact configuration.
Examples
7B model in FP16, 8K context (defaults)
Parameters 7B · FP16 · context 8192 · batch 1 · layers 32 · kvHeads 8 · headDim 128
Weights 14 GB + KV cache 1.07 GB = Total 15.07 GB
Fits: 16 GB cards (0.93 GB free), 24 GB, 48 GB, 80 GB. Does NOT fit 8 GB or 12 GB.
70B model in Q4_K_M, 4K context, Llama-2-70B shape
Parameters 70B · Q4_K_M (~0.61 B/weight) · context 4096 · layers 80
Weights 42.44 GB + KV cache 1.34 GB = Total 43.78 GB
Fits: 48 GB (RTX A6000 / L40S, 4.22 GB free) and 80 GB. Not a 24 GB RTX 4090.
8B model in INT4, no context needed
Parameters 8B · INT4 · context 0
Weights 4 GB + KV cache 0 GB = Total 4 GB → fits an 8 GB card with 4 GB to spare.
Good to know
- Units are decimal GB (10⁹ bytes), matching how “70B” and card sizes are quoted.
- KV formula: 2 × layers × context × kvHeads × headDim × 2 bytes (FP16 K+V) × batch. GQA models only cache the KV heads, so 8 ≠ 32 matters — check the model card.
- The total is a floor: activations, CUDA context, and inference-engine overhead (typically 0.5–2 GB) are not included — leave headroom.
- Private: all math runs locally in your browser; nothing is sent to a server.
- Related tools: LLM Cost Calculator, Token Estimator, Model Picker.