Skip to content

Model Quantization Explained

Cheatsheet of LLM quantization formats: bits per weight, the memory a 7B model needs in each, and the quality tradeoffs — FP32 to 4-bit and the GGUF naming.

Quantization shrinks model weights from 32-bit floats to 8, 4, or fewer bits. The prize is memory: a 7B model drops from ~26 GB to ~4 GB. The cost is quality — how much depends on the format and the model's size.

Reference table · 14 entries
14 of 14 rows
Full & half precision
FP3232~26 GBTraining ground truth; never needed for inference.
FP1616~13 GBDefault inference on GPU; lossless in practice.
BF1616~13 GBSame size as FP16, wider dynamic range — the training default.
8-bit
INT88~7 GBClassic integer quantization; solid quality.
Q8_08.5~7.5 GBGGUF simple per-block scaling; near-FP16 quality.
4-bit (the local sweet spot)
Q4_K_M~4.8~4.1 GBThe recommended default — k-quant, quality-aware mixes.
Q4_04.5~3.8 GBLegacy simple 4-bit; superseded by k-quants.
NF4 / GPTQ 44~3.5 GBQLoRA finetuning (NF4) / GPU serving (GPTQ, AWQ).
Aggressive
Q3_K_M~3.9~3.3 GBNoticeable degradation on small models; acceptable on 30B+.
Q2_K~3.4~2.8 GBLast resort — measurable quality loss on most tasks.
1.58-bit~2variesBitNet-era research territory, not general use.
Rules of thumb
bigger model, lower bits——A 13B at Q4 usually beats a 7B at Q8 — prefer size over precision.
KV cache——Quantizing weights doesn't shrink the cache; long contexts still need memory.
perplexity check——Validate a quant against the FP16 baseline before trusting it.