What GPU Do I Need for Local LLMs? RTX 5090 / RTX 4090 / M4 Max Estimator

What GPU Do I Need is a free hardware estimator that finds the cheapest GPU for your LLM — RTX 5090, RTX 4090, M4 Max, A100. No signup, instant results.

Free in-browser hardware estimator that recommends the cheapest GPU (NVIDIA RTX 5090, RTX 4090, RTX 3090, Apple Silicon M3/M4 Max, A100 80GB, H100 80GB) to run your chosen LLM at your target context length and quantization. Supports 63 open-weight models — Llama 4 Maverick/Scout, Gemma 4 12B/26B/31B, Kimi K2.7-Code, DeepSeek R1, Qwen 3, Phi-4, Mistral Small 3.2 and more.

Configuration

4K tokens
1K 128K
1x
1 32

Total VRAM Required

-- GB
Weights: 0.0 GB
KV Cache: 0.0 GB
Overhead: 0.0 GB

Recommended Hardware

Consumer
Apple Silicon
Server / Datacenter
System RAM Needed
0 GB

What GPU Do I Need for Local LLMs? A Hardware Planning Guide

Understanding LLM Hardware Requirements

When planning to run local Large Language Models (LLMs), the primary bottleneck is almost always memory (VRAM), not purely compute speed. If a model fits entirely within your GPU's VRAM, it will run incredibly fast. If it spills over into system RAM (often called "offloading" using tools like Llama.cpp), inference slows down drastically. This tool calculates the necessary VRAM layout to help you choose the right graphics card for your specific AI workload, preventing costly hardware mistakes.

The total formula for LLM VRAM consumption involves three components: the base model weights (scaled by quantization), the context window memory (KV Cache), and baseline operational overhead. By tweaking these fields in the estimator above, you can accurately map your requirements to consumer GPUs (like the RTX 4090 or RX 7900 XTX) or workstation silicon (like Apple's M-series Unified Memory).

Does Quantization Matter for GPU Choices?

Yes, immensely. Running an uncompressed 8-bit or 16-bit model requires massive datacenter infrastructure. Consumer setups thrive on quantization—techniques like GGUF (Q4_K_M) or AWQ/ExLlama formats compress models down to ~0.5 to 0.7 GB per billion parameters. This means a 7B to 9B model can comfortably fit on an inexpensive 8GB GPU. When sizing your graphics card, always estimate based on the quantization level you actually intend to run.

Apple Silicon vs Nvidia GeForce

Traditional discrete GPUs (Nvidia, AMD) have a hard wall: they top out at 24GB on consumer tiers (like the RTX 3090/4090). To run extremely massive models like Mixtral 8x22B or Llama 3 70B, you either need multiple graphics cards or workstation-class hardware. Apple Silicon (M1/M2/M3/M4 Max and Ultra) changes this dynamic by sharing memory between the CPU and GPU. A Mac Studio with 128GB of Unified Memory can run massive parameter sizes fluidly, effectively functioning as a high-VRAM GPU powerhouse without the associated PCI-e complexities.

Training vs. Inference on your GPU

Inference (generation) is highly predictable in memory usage. But if you plan to fine-tune an LLM using LoRA or QLoRA, memory depends as much on sequence length as on model size: every token of context needs activation memory (and the full-vocabulary logits), on top of the weights and the adapters' optimizer states. The training overlay uses the same estimate as the Can I Fine-Tune LLM calculator (Unsloth, gradient checkpointing). If you toggle the Training overlay in this estimator, you will immediately see why tuning typically pushes local developers out of the 8GB tier and into 16GB, 24GB, or multi-card setups.

Sources & Methodology

GPU recommendations and VRAM estimates are based on empirical testing, manufacturer specifications, and established performance tiers. Key references:

Recommendations represent general guidance. Actual performance depends on driver updates, cooling, thermal throttling, and workload specifics. Always benchmark your exact model + hardware combo before committing to a purchase.

Related Tools

🧮 Can I Run LLM? 🎯 Can I Fine-Tune LLM? 🏆 LLM Leaderboard 💻 Best LLM for Coding 🧠 Best LLM for Reasoning 📐 Best LLM for Math

Related Guides

📖 Best GPU for Running LLM Locally 📖 How to Calculate VRAM Requirements 📖 Best LLM for Indian Developers 2026 📖 Claude vs GPT vs Gemini 2026 📖 Best Free LLM API 2026 📖 LLM Benchmark Scores Explained 📖 How to Run LLM with Ollama