Every figure on this page is a real file size. 66 GGUF checkpoints across 11 models and 6 formats were measured by reading the x-linked-size header on each file's Hugging Face download URL, then divided by the parameter count the model's own safetensors index reports. No format delivers the bit width its name implies.
A GGUF file is not uniformly quantized. The K-quant formats apply the named block type to most attention and feed-forward weights, but token embeddings and the output projection are commonly held wider, and every quantized block carries scale and minimum values alongside its packed data. The name describes the dominant block. The file size is what you have to fit on the card.
| Format | Nominal bits | Measured mean | Min | Max | Over nominal | Size vs FP16 |
|---|---|---|---|---|---|---|
Q3_K_M | 3.0 | 4.003 | 3.888 | 4.271 | +33.4% | 25% |
IQ4_XS | 4.0 | 4.412 | 4.318 | 4.642 | +10.3% | 28% |
Q4_K_M | 4.0 | 4.914 | 4.826 | 5.110 | +22.8% | 31% |
Q5_K_M | 5.0 | 5.723 | 5.669 | 5.830 | +14.5% | 36% |
Q6_K | 6.0 | 6.571 | 6.564 | 6.596 | +9.5% | 41% |
Q8_0 | 8.0 | 8.508 | 8.502 | 8.533 | +6.3% | 53% |
The headline is Q4_K_M, the most widely distributed format on Hugging Face, at 4.914 bits per weight rather than 4. The relative overshoot is worst at the low end: Q3_K_M measures 4.003 bits against a nominal 3, a 33 percent overshoot, and in practice it lands almost exactly where a nominal 4-bit format is supposed to. If you are choosing Q3_K_M over Q4_K_M to save memory, the real saving is 19 percent, not the 25 percent the names suggest.
Stated against FP16, every format saves less than its name promises. The gap is the part that ends up on your GPU.
| Format | Implied saving | Measured saving | Shortfall |
|---|---|---|---|
Q3_K_M | 81% | 75% | 6 pt |
IQ4_XS | 75% | 72% | 3 pt |
Q4_K_M | 75% | 69% | 6 pt |
Q5_K_M | 69% | 64% | 5 pt |
Q6_K | 62% | 59% | 4 pt |
Q8_0 | 50% | 47% | 3 pt |
The spread across models is real and it is systematic. Small models overshoot most, because the tensors kept at higher precision, chiefly the embedding and output matrices, are a larger fraction of a smaller parameter count. A 1.5B model and a 32B model quantized to the same nominal format do not land on the same bits per weight.
| Model | Parameters | Q3_K_M | IQ4_XS | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|---|---|---|
| Qwen2.5-1.5B-Instruct | 1.54B | 4.271 | 4.642 | 5.110 | 5.830 | 6.596 | 8.533 |
| Qwen2.5-3B-Instruct | 3.09B | 4.123 | 4.508 | 5.003 | 5.768 | 6.580 | 8.517 |
| Mistral-7B-Instruct-v0.3 | 7.25B | 3.888 | 4.318 | 4.826 | 5.669 | 6.564 | 8.502 |
| Qwen2.5-7B-Instruct | 7.62B | 4.001 | 4.431 | 4.919 | 5.720 | 6.570 | 8.507 |
| DeepSeek-R1-Distill-Qwen-7B | 7.62B | 4.001 | 4.431 | 4.919 | 5.720 | 6.570 | 8.507 |
| Meta-Llama-3.1-8B-Instruct | 8.03B | 4.004 | 4.431 | 4.902 | 5.711 | 6.571 | 8.509 |
| phi-4 | 14.66B | 4.018 | 4.334 | 4.940 | 5.787 | 6.565 | 8.503 |
| Qwen2.5-14B-Instruct | 14.77B | 3.975 | 4.398 | 4.868 | 5.692 | 6.567 | 8.505 |
| DeepSeek-R1-Distill-Qwen-14B | 14.77B | 3.975 | 4.398 | 4.868 | 5.692 | 6.567 | 8.505 |
| DeepSeek-R1-Distill-Qwen-32B | 32.76B | 3.891 | 4.320 | 4.847 | 5.680 | 6.565 | 8.502 |
| Qwen2.5-32B-Instruct | 32.76B | 3.891 | 4.320 | 4.847 | 5.680 | 6.565 | 8.502 |
The same measurements expressed as the number you compare against card capacity. A 24 GB card reports 22.35 GiB, so read these against that rather than against the advertised figure.
| Model | Q3_K_M | IQ4_XS | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|---|---|
| Qwen2.5-1.5B-Instruct | 0.77 | 0.83 | 0.92 | 1.05 | 1.19 | 1.53 |
| Qwen2.5-3B-Instruct | 1.48 | 1.62 | 1.80 | 2.07 | 2.36 | 3.06 |
| Mistral-7B-Instruct-v0.3 | 3.28 | 3.64 | 4.07 | 4.78 | 5.54 | 7.17 |
| Qwen2.5-7B-Instruct | 3.55 | 3.93 | 4.36 | 5.07 | 5.82 | 7.54 |
| DeepSeek-R1-Distill-Qwen-7B | 3.55 | 3.93 | 4.36 | 5.07 | 5.82 | 7.54 |
| Meta-Llama-3.1-8B-Instruct | 3.74 | 4.14 | 4.58 | 5.34 | 6.14 | 7.95 |
| phi-4 | 6.86 | 7.40 | 8.43 | 9.88 | 11.20 | 14.51 |
| Qwen2.5-14B-Instruct | 6.84 | 7.56 | 8.37 | 9.79 | 11.29 | 14.62 |
| DeepSeek-R1-Distill-Qwen-14B | 6.84 | 7.56 | 8.37 | 9.79 | 11.29 | 14.62 |
| DeepSeek-R1-Distill-Qwen-32B | 14.84 | 16.48 | 18.49 | 21.66 | 25.04 | 32.43 |
| Qwen2.5-32B-Instruct | 14.84 | 16.48 | 18.49 | 21.66 | 25.04 | 32.43 |
Two rows are worth dwelling on. Qwen2.5-32B-Instruct at Q4_K_M is 18.49 GiB. Sized at a nominal 4 bits you would predict 15.25 GiB and conclude it leaves plenty of room on a 24 GB card. It does not: after the KV cache and runtime overhead it does not fit at any useful context length. To check that for your own combination, use the GPU fit calculator, which uses these measured figures directly.
File sizes come from an HTTP HEAD request against each file's resolve/main URL on Hugging Face, reading the x-linked-size response header. That header reports the size of the object the LFS pointer resolves to, so it is the real download size without transferring the file. Parameter counts come from the safetensors.total field of the Hugging Face model API for the corresponding unquantized repository, which is derived from the tensor shapes in the checkpoint index rather than from the model's name.
Effective bits per weight is then file bytes times 8, divided by parameter count. All measurements were taken on 22 August 2026 from the bartowski GGUF repositories, which publish a consistent quantization matrix across models and are the most downloaded source for these formats. Quantization is deterministic given the same source checkpoint and llama.cpp version, so these figures are stable, but a repository requantized with a newer llama.cpp can shift slightly.
One caveat worth stating plainly. The models measured here are dense transformers in the 1.5B to 32B range. Mixture-of-experts checkpoints distribute parameters differently and were not sampled, so do not carry these ratios across to them. Vision and audio towers were also excluded.