Model Quantization for Normal People
GGUF, Q4_K_M, Q8_0, BF16: the quantization naming is designed to confuse. Here's what it actually means and how to pick the right format for your hardware.
Quantization lets you run a 70B parameter model on hardware that was never designed to handle a 70B parameter model. The tradeoff is quality. The trick is knowing how much quality you're trading for how much hardware headroom.
Here's the answer up front: quantization is the process of reducing the numerical precision of a model's weights to make it smaller and faster to run. A full-precision model stores each weight as a 32-bit float. A quantized model stores each weight in a compressed format: 8 bits, 4 bits, sometimes less. Smaller weights mean smaller model files, lower VRAM requirements, and faster inference. The tradeoff is a small but measurable reduction in output quality that varies by how aggressively you quantize.
If you're running models locally, you need to understand this. If you're not running models locally, skip this one and come back when you are.
the naming explained
BF16 / FP16: Near full-quality. 16-bit precision. Requires the most VRAM: roughly 2GB per billion parameters. A 7B model in BF16 needs ~14GB VRAM. Good if you have the hardware. Most people running on consumer hardware don't.
Q8_0: 8-bit quantization. About 1GB per billion parameters. Very small quality loss compared to full precision, often imperceptible in practice. A 7B model needs ~7GB VRAM. A good choice when you have enough VRAM and want near-full quality.
Q4_K_M: 4-bit quantization, K-quant method, medium variant. About 0.5GB per billion parameters. A 7B model needs ~4GB VRAM. Noticeable quality reduction versus Q8 in careful comparison, but still very usable for most practical tasks. This is the most common "I'm running on a normal GPU" format.
Q4_K_S: 4-bit, small variant. Smaller than Q4_K_M, slightly lower quality. Use when you need to squeeze into a smaller VRAM budget.
Q2_K: 2-bit quantization. Very small, significant quality loss. Only worth considering if you genuinely can't fit anything bigger. The outputs are noticeably degraded.
The K-quant formats (the ones with K in the name) use a more sophisticated quantization method that preserves more of the model's quality than the older Q-formats. For the same bit depth, K-quant is better. Q4_K_M is better than Q4_0.
GGUF: the file format
GGUF is the container format used by llama.cpp and Ollama to store quantized models. If you're downloading a model for local use, you're almost certainly working with GGUF files. The format encodes both the model weights and the metadata needed to run them.
When you pull a model in Ollama, it's handling the GGUF under the hood. When you download a model directly from Hugging Face for llama.cpp, you're picking the GGUF variant.
how to pick
Have 16GB+ VRAM (high-end consumer or prosumer GPU): Run Q8_0 for 7B/13B models, Q4_K_M for 70B models.
Have 8-12GB VRAM (RTX 3080/4070 class): Q4_K_M for 7B/13B models. 70B models are marginal. You'll need to offload layers to CPU, which slows things down significantly.
Have 6-8GB VRAM (RTX 3060/4060 class): Q4_K_M or Q4_K_S for 7B models. Larger models run too slowly to be useful.
CPU only / no discrete GPU: Q4_K_M still works via llama.cpp with CPU inference. Expect slower generation: 1-3 tokens/second on a modern CPU for 7B models. Usable for non-interactive tasks; frustrating for conversation.
Apple Silicon (M-series Mac): Unified memory means you can run larger models than discrete GPU comparisons suggest. An M3 Pro with 36GB can comfortably run Q8_0 for 13B models and Q4_K_M for 70B. Ollama handles this well natively.
quality loss in practice
The real-world quality difference between Q8_0 and Q4_K_M on most practical tasks is smaller than benchmarks suggest. For conversation, summarization, code generation, and classification, Q4_K_M is usually "good enough." The quality loss is most visible on tasks requiring precise factual recall and complex multi-step reasoning, exactly the tasks where you should probably be using retrieval grounding anyway.
Test on your actual use case. Don't rely on benchmark comparisons.
Try it today
| Step | What you do | Why it pays off |
|---|---|---|
| 1. Check your VRAM | nvidia-smi on Linux/Windows or Activity Monitor on Mac to see available VRAM | Determines which quantization sizes you can actually run |
| 2. Start with Q4_K_M | Pull the Q4_K_M variant of your target model via Ollama or direct GGUF download | Best practical tradeoff for most hardware, and your baseline for comparison |
| 3. Test against Q8_0 on your actual tasks | Run both variants on 10-15 real examples from your use case. Compare outputs. | Quality difference varies by task. Know if Q8 is worth the VRAM premium for YOUR use case. |
The bottom line
Quantization lets you run large models on hardware that wasn't built for them. Q4_K_M is the practical sweet spot for most consumer hardware. Q8_0 when you have the VRAM and want near-full quality. Anything below Q4 only when you have no other choice.
The naming is intimidating until you understand it. Then it's just arithmetic.
Dru Edwards