Local AI Hardware
LLM Quantization Explained: Q4 vs Q8 — What's the Difference and Which Should You Use?
Last updated: 2026-03-12
If you have downloaded local models, you have seen names like Q4_K_M, Q8_0, F16, and F32. Those labels decide whether a model fits your machine, how much quality you keep, and how fast generation feels. This guide explains quantization in practical terms so you can choose the right format by hardware tier instead of guessing from forum comments.
🛠️ Developer's Note
Quantization is probably the most misunderstood topic in local AI. I see people either avoiding it entirely (and running out of VRAM) or using maximum compression without understanding the quality trade-off. This guide exists to make that trade-off concrete and quantifiable.
What Is Quantization?
Quantization reduces weight precision so model files use less memory. Full precision F32 stores each weight in 32 bits (4 bytes). F16 stores 16 bits. Q8 uses roughly 8 bits. Q4 uses roughly 4 bits. Lower precision means smaller model footprints, which directly improves the chance of fitting a model inside available VRAM or unified memory.
The trade-off is output fidelity. As precision drops, subtle reasoning quality can decline. For most local usage, the gain from being able to run a better-sized model usually outweighs tiny per-token quality changes from higher precision.
Quantization Formats Compared
| Format | Bits/Weight | VRAM vs F16 | Quality Loss | Use Case |
|---|---|---|---|---|
| F32 | 32-bit | 2x larger than F16 | None | Training/debug only |
| F16 | 16-bit | Baseline | None | High-memory inference |
| Q8_0 | 8-bit | ~50% smaller | Very low (~1%) | Best quality/VRAM balance |
| Q4_K_M | 4-bit mixed | ~70% smaller | Low-medium (~3-5%) | Consumer GPU sweet spot |
| Q4_K_S | 4-bit small | Slightly smaller than Q4_K_M | Similar | Very tight memory |
| Q2_K | 2-bit | ~85% smaller | Noticeable | Last-resort fit |
Practical Example: Llama 3.1 8B
Quantization impact is easiest to see on one model. A typical 8B setup roughly looks like this:
- F16: ~16GB
- Q8_0: ~9GB
- Q4_K_M: ~5GB
- Q2_K: ~3GB
On an 8GB GPU, F16 is impossible, Q8 is borderline or non-fit, and Q4 is practical. That single decision determines whether you can actually use the model daily.
Which Quantization Should You Pick?
Use memory-first decision rules. They are simple and usually correct.
- 8GB VRAM: default to Q4_K_M.
- 12-16GB VRAM: choose Q8_0 for small models, Q4_K_M for bigger ones.
- 24GB+ VRAM: Q8_0 for most workloads; F16 for smaller models up to ~13B.
- Apple Silicon: Q4_K_M by default; move to Q8_0 on 32GB+ unified memory.
If uncertain, benchmark both Q4_K_M and Q8_0 on the same prompt set and pick the highest quality that still gives stable speed.
Where to Find Quantized Models
Most local users discover quantized GGUF builds through Hugging Face publishers and Ollama model registry. Historical sources such as TheBloke archives are still useful, while newer maintainers like bartowski and unsloth provide updated releases for recent base models.
Ollama usually abstracts format complexity, but understanding quantization helps when you choose model tags and evaluate fit/performance trade-offs.
Check VRAM Before You Download
Use Can I Run LLM to test model-size and quantization combinations against your hardware. It is faster than trial pulls and avoids disk and time waste.
Check your GPU before you download
Can I Run LLM — Free VRAM Calculator →Quantization is not a niche optimization; it is the core lever that makes local LLMs usable on consumer hardware. Once you map format choice to memory tier, model selection becomes predictable, and local AI setup stops feeling random.
How Quantization Affects Real Tasks
Quality differences are easiest to notice in edge cases: long chained reasoning, strict JSON output, and nuanced code generation. For many day-to-day tasks such as summarization, rewriting, or quick Q&A, Q4_K_M can be very close to Q8_0 while using much less memory. That is why Q4 remains the default for consumer setups.
If your workflow demands high precision outputs, test Q8_0 on the same prompt set and measure error rates rather than trusting generic \"quality score\" claims. In many cases, prompt engineering and better context formatting improve output more than changing quantization alone.
Benchmark Method You Can Reuse
Create a small benchmark pack with 20 prompts covering your real use cases: one coding task, one data extraction task, one long-summary task, one reasoning chain, and one strict format output task. Run the same pack on Q4 and Q8 with temperature fixed, then compare three metrics: latency, format correctness, and subjective quality.
- Measure first-token latency and full response time separately.
- Count parse failures for JSON/tool-output prompts.
- Track hallucinated facts in multi-step reasoning prompts.
- Keep context length fixed so results are comparable.
This method gives a practical, repeatable way to choose quantization for your team instead of relying on one-off impressions.
Common Quantization Myths
- Myth: \"Q8 is always better.\" Reality: only if it still fits and keeps speed acceptable.
- Myth: \"Q4 is unusable for serious work.\" Reality: many production local workflows run on Q4.
- Myth: \"Lower bits always means huge quality loss.\" Reality: architecture and prompt quality matter too.
The best quantization is the one that keeps your model reliable in your real workflow. If Q8 slows your loop so much that you stop using the model, Q4 with better prompts will often deliver more practical value. Optimize for repeatable productivity, not format prestige.
In short, pick the highest precision that still preserves fast iteration on your machine, because iteration speed is usually the strongest predictor of long-term local AI adoption.
Limitations / When NOT to Use This
- Quality impact varies by model architecture — some models (especially smaller ones under 7B) degrade more noticeably at Q4 than larger models do at the same quantization level
- The Q4 vs Q8 comparison here focuses on GGUF format; other formats (AWQ, GPTQ, EXL2) have different quality/size trade-offs not covered in this guide
- Benchmark scores for quantized models can be misleading — perplexity measurements may show small differences that become noticeable in real conversation quality, especially for creative writing or nuanced reasoning
- Mixing quantization levels within a model (e.g., K-quant variants) adds complexity not fully captured by simple Q4/Q8 labels