What GPU Do I Need for Local LLMs? RTX 5090 / RTX 4090 / M4 Max Estimator
What GPU Do I Need is a free hardware estimator that finds the cheapest GPU for your LLM — RTX 5090, RTX 4090, M4 Max, A100. No signup, instant results.
Free in-browser hardware estimator that recommends the cheapest GPU (NVIDIA RTX 5090, RTX 4090, RTX 3090, Apple Silicon M3/M4 Max, A100 80GB, H100 80GB) to run your chosen LLM at your target context length and quantization. Supports 63 open-weight models — Llama 4 Maverick/Scout, Gemma 4 12B/26B/31B, Kimi K2.7-Code, DeepSeek R1, Qwen 3, Phi-4, Mistral Small 3.2 and more.
Configuration
Total VRAM Required
Recommended Hardware
What GPU Do I Need for Local LLMs? A Hardware Planning Guide
Understanding LLM Hardware Requirements
When planning to run local Large Language Models (LLMs), the primary bottleneck is almost always memory (VRAM), not purely compute speed. If a model fits entirely within your GPU's VRAM, it will run incredibly fast. If it spills over into system RAM (often called "offloading" using tools like Llama.cpp), inference slows down drastically. This tool calculates the necessary VRAM layout to help you choose the right graphics card for your specific AI workload, preventing costly hardware mistakes.
The total formula for LLM VRAM consumption involves three components: the base model weights (scaled by quantization), the context window memory (KV Cache), and baseline operational overhead. By tweaking these fields in the estimator above, you can accurately map your requirements to consumer GPUs (like the RTX 4090 or RX 7900 XTX) or workstation silicon (like Apple's M-series Unified Memory).
Does Quantization Matter for GPU Choices?
Yes, immensely. Running an uncompressed 8-bit or 16-bit model requires massive datacenter infrastructure. Consumer setups thrive on quantization—techniques like GGUF (Q4_K_M) or AWQ/ExLlama formats compress models down to ~0.5 to 0.7 GB per billion parameters. This means a 7B to 9B model can comfortably fit on an inexpensive 8GB GPU. When sizing your graphics card, always estimate based on the quantization level you actually intend to run.
Apple Silicon vs Nvidia GeForce
Traditional discrete GPUs (Nvidia, AMD) have a hard wall: they top out at 24GB on consumer tiers (like the RTX 3090/4090). To run extremely massive models like Mixtral 8x22B or Llama 3 70B, you either need multiple graphics cards or workstation-class hardware. Apple Silicon (M1/M2/M3/M4 Max and Ultra) changes this dynamic by sharing memory between the CPU and GPU. A Mac Studio with 128GB of Unified Memory can run massive parameter sizes fluidly, effectively functioning as a high-VRAM GPU powerhouse without the associated PCI-e complexities.
Training vs. Inference on your GPU
Inference (generation) is highly predictable in memory usage. But if you plan to fine-tune an LLM using LoRA or QLoRA, memory depends as much on sequence length as on model size: every token of context needs activation memory (and the full-vocabulary logits), on top of the weights and the adapters' optimizer states. The training overlay uses the same estimate as the Can I Fine-Tune LLM calculator (Unsloth, gradient checkpointing). If you toggle the Training overlay in this estimator, you will immediately see why tuning typically pushes local developers out of the 8GB tier and into 16GB, 24GB, or multi-card setups.
Sources & Methodology
GPU recommendations and VRAM estimates are based on empirical testing, manufacturer specifications, and established performance tiers. Key references:
- llama.cpp GitHub — GPU backend compatibility and VRAM allocation
- Tim Dettmers: Which GPU for Deep Learning — consumer GPU performance tier analysis
- Ollama Hardware Requirements — recommended hardware per model
- Apple M-series Unified Memory performance — Metal GPU compute specs
Recommendations represent general guidance. Actual performance depends on driver updates, cooling, thermal throttling, and workload specifics. Always benchmark your exact model + hardware combo before committing to a purchase.