Hardware Guide
Best Local LLM for 8 GB, 12 GB, 16 GB and 24 GB VRAM in 2026
Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset
The first question with a local model is simple: will it fit? This page lists the largest models that fit in each common VRAM size, using the same memory engine as the id8 calculators. Figures are for Q4_K_M quantization with an 8,192-token context, a sensible everyday setup.
{{localCount}} models, every quantization level, any context length.
Check your exact GPU and model →How these tables are calculated
Every figure comes from the model's own architecture file (layers, hidden size, attention heads, vocabulary) through the memory engine that powers Can I Run LLM?. The total includes the quantized weights, the cache for an 8,192-token context, and runtime overhead. A model is listed under a VRAM size only if it fits with about 5% to spare.
Each table shows the largest models that fit, biggest first. Bigger is not always better for your task, but within a model family more parameters generally means better answers.
8 GB VRAM
Typical cards: RTX 4060, RTX 3060 Ti, RTX 3070, RTX 5060.
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Qwen3.5 9B | Alibaba | 9.65B | 6.7 GB | — |
| Ministral 8B | Mistral AI | 8.9B | 6.9 GB | ministral-3:8b |
| ZAYA1-8B | Zyphra | 8.84B | 6.1 GB | — |
| Llama 3.1 8B | Meta | 8B | 6.3 GB | llama3.1:8b |
| Qwen 3 8B | Alibaba | 8B | 6.5 GB | qwen3:8b |
| DeepSeek R1 Distill 8B | DeepSeek | 8B | 6.3 GB | deepseek-r1:8b |
| Qwen 2.5 7B | Alibaba | 7.62B | 5.6 GB | qwen2.5:7b |
| Qwen 2.5-Coder 7B | Alibaba | 7.62B | 5.6 GB | qwen2.5-coder:7b |
At 8 GB you are choosing among small models. They handle chat, summarising and simple coding help well, and are weaker on long reasoning.
12 GB VRAM
Typical cards: RTX 3060 12GB, RTX 4070, RTX 5070.
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Qwen 2.5 14B | Alibaba | 14.77B | 10.9 GB | qwen2.5:14b |
| Qwen 3 14B | Alibaba | 14.77B | 10.6 GB | qwen3:14b |
| DeepSeek R1 Distill 14B | DeepSeek | 14.77B | 10.9 GB | deepseek-r1:14b |
| Phi-4 14B | Microsoft | 14.66B | 10.9 GB | phi4:14b |
| Phi-4-reasoning-plus 14B | Microsoft | 14.66B | 10.9 GB | phi4-reasoning:plus |
| Phi-3 Medium 14B | Microsoft | 14B | 10.5 GB | phi3:14b |
| Ministral 14B | Mistral AI | 13.9B | 10.1 GB | ministral-3:14b |
| Gemma 3 12B | 12B | 8.6 GB | gemma3:12b |
12 GB is the first size where mid-size models become comfortable, with room for a longer context.
16 GB VRAM
Typical cards: RTX 4060 Ti 16GB, RTX 4080, RTX 5060 Ti 16GB, RTX 5070 Ti, RTX 5080.
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Qwen 2.5 14B | Alibaba | 14.77B | 10.9 GB | qwen2.5:14b |
| Qwen 3 14B | Alibaba | 14.77B | 10.6 GB | qwen3:14b |
| DeepSeek R1 Distill 14B | DeepSeek | 14.77B | 10.9 GB | deepseek-r1:14b |
| Phi-4 14B | Microsoft | 14.66B | 10.9 GB | phi4:14b |
| Phi-4-reasoning-plus 14B | Microsoft | 14.66B | 10.9 GB | phi4-reasoning:plus |
| Phi-3 Medium 14B | Microsoft | 14B | 10.5 GB | phi3:14b |
| Ministral 14B | Mistral AI | 13.9B | 10.1 GB | ministral-3:14b |
| Gemma 3 12B | 12B | 8.6 GB | gemma3:12b |
16 GB is a good balance of price and capability for a personal machine.
24 GB VRAM
Typical cards: RTX 3090, RTX 4090.
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Qwen3.6 35B-A3B | Alibaba | 35B | 21.2 GB | qwen3.6:35b |
| QwQ 32B | Alibaba | 32.5B | 21.5 GB | qwq:32b |
| Qwen 2.5 32B | Alibaba | 32B | 21.2 GB | qwen2.5:32b |
| Qwen 2.5-Coder 32B | Alibaba | 32B | 21.2 GB | qwen2.5-coder:32b |
| Qwen 3 32B | Alibaba | 32B | 21.2 GB | qwen3:32b |
| DeepSeek R1 Distill 32B | DeepSeek | 32B | 21.2 GB | deepseek-r1:32b |
| Gemma 4 31B | 30.7B | 20.5 GB | gemma4:31b | |
| Qwen 3 30B-A3B | Alibaba | 30B | 18.8 GB | qwen3:30b-a3b |
24 GB runs the 30-billion-parameter class at Q4, which is where local models start to feel close to mid-tier API models for many tasks.
32 GB VRAM
Typical card: RTX 5090.
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Mixtral 8x7B | Mistral AI | 46.7B | 28.7 GB | mixtral:8x7b |
| Qwen3.6 35B-A3B | Alibaba | 35B | 21.2 GB | qwen3.6:35b |
| QwQ 32B | Alibaba | 32.5B | 21.5 GB | qwq:32b |
| Qwen 2.5 32B | Alibaba | 32B | 21.2 GB | qwen2.5:32b |
| Qwen 2.5-Coder 32B | Alibaba | 32B | 21.2 GB | qwen2.5-coder:32b |
| Qwen 3 32B | Alibaba | 32B | 21.2 GB | qwen3:32b |
If the model you want does not fit
- Use a smaller quantization. Q3 needs less memory than Q4, with a larger quality loss.
- Shorten the context. Going from 8K to 4K tokens frees memory.
- Pick the next size down in the same family.
- Offload some layers to RAM. It works, at a large cost in speed.
Here is how quantization and context change the memory for one mid-size model, Qwen 3 14B:
| Quantization | VRAM at 4K context | VRAM at 8K context | VRAM at 32K context |
|---|---|---|---|
| Q3_K_M | 8.4 GB | 9.0 GB | 12.8 GB |
| Q4_K_M | 10.0 GB | 10.6 GB | 14.4 GB |
| Q5_K_M | 11.4 GB | 12.1 GB | 15.8 GB |
| Q6_K | 12.9 GB | 13.5 GB | 17.3 GB |
| Q8_0 | 16.2 GB | 16.9 GB | 20.6 GB |
Apple Silicon
Macs share memory between the system and the GPU. Roughly three quarters of the total can be used for the model, so a 32 GB Mac behaves like a 24 GB card and a 64 GB Mac like a 48 GB one. See Run LLM on Apple Silicon Mac.
Choosing within a size
- For coding help: prefer a model with a published coding score; see Best LLM for Coding and switch on the self-host filter.
- For reasoning: models with a thinking mode do better and are slower.
- For images: choose a model marked as multimodal.
- For speed: a smaller model that fits with room to spare is faster than a larger one that barely fits.
Check your own setup
These tables use one setting. Your GPU, quantization and context length will differ. Can I Run LLM? calculates the exact figure for any combination, and What GPU Do I Need? works the other way: pick a model and it tells you the cheapest hardware that runs it.
Frequently asked questions
What is the best local LLM for 8 GB VRAM?
The largest models that fit in 8 GB at Q4_K_M with an 8K context are listed in the 8 GB table on this page. Models of roughly 7 to 9 billion parameters are the practical ceiling.
What can a 24 GB card like the RTX 4090 run?
At Q4_K_M with an 8K context, a 24 GB card runs models up to roughly the 30-billion-parameter class. The 24 GB table lists the current options.
What does Q4_K_M mean?
It is a 4-bit quantization of the model's weights in the GGUF format. It cuts memory to roughly a quarter of full precision with a small loss in quality, and is the usual default for local use.
Why does context length change the VRAM needed?
The model keeps a cache for every token in the context. A longer context means a larger cache, on top of the weights.
Can I run a model that does not fit in VRAM?
Yes, by placing some layers in system RAM, but it becomes much slower. For comfortable speed the whole model should fit in VRAM.