| Full & half precision | |||
|---|---|---|---|
| FP32 | 32 | ~26 GB | Training ground truth; never needed for inference. |
| FP16 | 16 | ~13 GB | Default inference on GPU; lossless in practice. |
| BF16 | 16 | ~13 GB | Same size as FP16, wider dynamic range — the training default. |
| 8-bit | |||
| INT8 | 8 | ~7 GB | Classic integer quantization; solid quality. |
| Q8_0 | 8.5 | ~7.5 GB | GGUF simple per-block scaling; near-FP16 quality. |
| 4-bit (the local sweet spot) | |||
| Q4_K_M | ~4.8 | ~4.1 GB | The recommended default — k-quant, quality-aware mixes. |
| Q4_0 | 4.5 | ~3.8 GB | Legacy simple 4-bit; superseded by k-quants. |
| NF4 / GPTQ 4 | 4 | ~3.5 GB | QLoRA finetuning (NF4) / GPU serving (GPTQ, AWQ). |
| Aggressive | |||
| Q3_K_M | ~3.9 | ~3.3 GB | Noticeable degradation on small models; acceptable on 30B+. |
| Q2_K | ~3.4 | ~2.8 GB | Last resort — measurable quality loss on most tasks. |
| 1.58-bit | ~2 | varies | BitNet-era research territory, not general use. |
| Rules of thumb | |||
| bigger model, lower bits | — | — | A 13B at Q4 usually beats a 7B at Q8 — prefer size over precision. |
| KV cache | — | — | Quantizing weights doesn't shrink the cache; long contexts still need memory. |
| perplexity check | — | — | Validate a quant against the FP16 baseline before trusting it. |