Hardware Guide

Best Local LLM for 8 GB, 12 GB, 16 GB and 24 GB VRAM in 2026

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

The first question with a local model is simple: will it fit? This page lists the largest models that fit in each common VRAM size, using the same memory engine as the id8 calculators. Figures are for Q4_K_M quantization with an 8,192-token context, a sensible everyday setup.

{{localCount}} models, every quantization level, any context length.

Check your exact GPU and model →

How these tables are calculated

Every figure comes from the model's own architecture file (layers, hidden size, attention heads, vocabulary) through the memory engine that powers Can I Run LLM?. The total includes the quantized weights, the cache for an 8,192-token context, and runtime overhead. A model is listed under a VRAM size only if it fits with about 5% to spare.

Each table shows the largest models that fit, biggest first. Bigger is not always better for your task, but within a model family more parameters generally means better answers.

8 GB VRAM

Typical cards: RTX 4060, RTX 3060 Ti, RTX 3070, RTX 5060.

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Qwen3.5 9BAlibaba9.65B6.7 GB—
Ministral 8BMistral AI8.9B6.9 GBministral-3:8b
ZAYA1-8BZyphra8.84B6.1 GB—
Llama 3.1 8BMeta8B6.3 GBllama3.1:8b
Qwen 3 8BAlibaba8B6.5 GBqwen3:8b
DeepSeek R1 Distill 8BDeepSeek8B6.3 GBdeepseek-r1:8b
Qwen 2.5 7BAlibaba7.62B5.6 GBqwen2.5:7b
Qwen 2.5-Coder 7BAlibaba7.62B5.6 GBqwen2.5-coder:7b

At 8 GB you are choosing among small models. They handle chat, summarising and simple coding help well, and are weaker on long reasoning.

12 GB VRAM

Typical cards: RTX 3060 12GB, RTX 4070, RTX 5070.

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Qwen 2.5 14BAlibaba14.77B10.9 GBqwen2.5:14b
Qwen 3 14BAlibaba14.77B10.6 GBqwen3:14b
DeepSeek R1 Distill 14BDeepSeek14.77B10.9 GBdeepseek-r1:14b
Phi-4 14BMicrosoft14.66B10.9 GBphi4:14b
Phi-4-reasoning-plus 14BMicrosoft14.66B10.9 GBphi4-reasoning:plus
Phi-3 Medium 14BMicrosoft14B10.5 GBphi3:14b
Ministral 14BMistral AI13.9B10.1 GBministral-3:14b
Gemma 3 12BGoogle12B8.6 GBgemma3:12b

12 GB is the first size where mid-size models become comfortable, with room for a longer context.

16 GB VRAM

Typical cards: RTX 4060 Ti 16GB, RTX 4080, RTX 5060 Ti 16GB, RTX 5070 Ti, RTX 5080.

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Qwen 2.5 14BAlibaba14.77B10.9 GBqwen2.5:14b
Qwen 3 14BAlibaba14.77B10.6 GBqwen3:14b
DeepSeek R1 Distill 14BDeepSeek14.77B10.9 GBdeepseek-r1:14b
Phi-4 14BMicrosoft14.66B10.9 GBphi4:14b
Phi-4-reasoning-plus 14BMicrosoft14.66B10.9 GBphi4-reasoning:plus
Phi-3 Medium 14BMicrosoft14B10.5 GBphi3:14b
Ministral 14BMistral AI13.9B10.1 GBministral-3:14b
Gemma 3 12BGoogle12B8.6 GBgemma3:12b

16 GB is a good balance of price and capability for a personal machine.

24 GB VRAM

Typical cards: RTX 3090, RTX 4090.

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Qwen3.6 35B-A3BAlibaba35B21.2 GBqwen3.6:35b
QwQ 32BAlibaba32.5B21.5 GBqwq:32b
Qwen 2.5 32BAlibaba32B21.2 GBqwen2.5:32b
Qwen 2.5-Coder 32BAlibaba32B21.2 GBqwen2.5-coder:32b
Qwen 3 32BAlibaba32B21.2 GBqwen3:32b
DeepSeek R1 Distill 32BDeepSeek32B21.2 GBdeepseek-r1:32b
Gemma 4 31BGoogle30.7B20.5 GBgemma4:31b
Qwen 3 30B-A3BAlibaba30B18.8 GBqwen3:30b-a3b

24 GB runs the 30-billion-parameter class at Q4, which is where local models start to feel close to mid-tier API models for many tasks.

32 GB VRAM

Typical card: RTX 5090.

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Mixtral 8x7BMistral AI46.7B28.7 GBmixtral:8x7b
Qwen3.6 35B-A3BAlibaba35B21.2 GBqwen3.6:35b
QwQ 32BAlibaba32.5B21.5 GBqwq:32b
Qwen 2.5 32BAlibaba32B21.2 GBqwen2.5:32b
Qwen 2.5-Coder 32BAlibaba32B21.2 GBqwen2.5-coder:32b
Qwen 3 32BAlibaba32B21.2 GBqwen3:32b

If the model you want does not fit

  1. Use a smaller quantization. Q3 needs less memory than Q4, with a larger quality loss.
  2. Shorten the context. Going from 8K to 4K tokens frees memory.
  3. Pick the next size down in the same family.
  4. Offload some layers to RAM. It works, at a large cost in speed.

Here is how quantization and context change the memory for one mid-size model, Qwen 3 14B:

QuantizationVRAM at 4K contextVRAM at 8K contextVRAM at 32K context
Q3_K_M8.4 GB9.0 GB12.8 GB
Q4_K_M10.0 GB10.6 GB14.4 GB
Q5_K_M11.4 GB12.1 GB15.8 GB
Q6_K12.9 GB13.5 GB17.3 GB
Q8_016.2 GB16.9 GB20.6 GB

Apple Silicon

Macs share memory between the system and the GPU. Roughly three quarters of the total can be used for the model, so a 32 GB Mac behaves like a 24 GB card and a 64 GB Mac like a 48 GB one. See Run LLM on Apple Silicon Mac.

Choosing within a size

  • For coding help: prefer a model with a published coding score; see Best LLM for Coding and switch on the self-host filter.
  • For reasoning: models with a thinking mode do better and are slower.
  • For images: choose a model marked as multimodal.
  • For speed: a smaller model that fits with room to spare is faster than a larger one that barely fits.

Check your own setup

These tables use one setting. Your GPU, quantization and context length will differ. Can I Run LLM? calculates the exact figure for any combination, and What GPU Do I Need? works the other way: pick a model and it tells you the cheapest hardware that runs it.

Frequently asked questions

What is the best local LLM for 8 GB VRAM?

The largest models that fit in 8 GB at Q4_K_M with an 8K context are listed in the 8 GB table on this page. Models of roughly 7 to 9 billion parameters are the practical ceiling.

What can a 24 GB card like the RTX 4090 run?

At Q4_K_M with an 8K context, a 24 GB card runs models up to roughly the 30-billion-parameter class. The 24 GB table lists the current options.

What does Q4_K_M mean?

It is a 4-bit quantization of the model's weights in the GGUF format. It cuts memory to roughly a quarter of full precision with a small loss in quality, and is the usual default for local use.

Why does context length change the VRAM needed?

The model keeps a cache for every token in the context. A longer context means a larger cache, on top of the weights.

Can I run a model that does not fit in VRAM?

Yes, by placing some layers in system RAM, but it becomes much slower. For comfortable speed the whole model should fit in VRAM.

Related