Can I Run LLM? Free Local LLM VRAM Calculator

CanIRunLLM is a free VRAM calculator that checks if your GPU can run 95 self-hosted LLMs — Llama 4, Gemma 4, Kimi K2, DeepSeek R1. No upload, no signup.

Use the id8 CanIRunLLM calculator to check if your GPU VRAM is sufficient for local AI models in 2026 — Llama 4 Maverick, Llama 4 Scout, Gemma 4 (12B, 26B, 31B), Kimi K2.7-Code, DeepSeek R1, Mistral Small 3.2, Phi-4, Qwen 3, Nemotron 3 Ultra, MiniMax M2.7 and 50 more open-weight models.

Free in-browser VRAM calculator. No upload, no signup, no telemetry. Updated October 2026.

Your Hardware

12 GB
4GB 192GB
32 GB
4GB 256GB

* Used for "Offloading" when model doesn't fit in VRAM. Slower than GPU only.

4096 Tokens
4K 128K+

* Longer context requires more VRAM for the KV cache.

Runs Great
<60% VRAM
Runs Well
60–80% VRAM
Tight Fit
80–95% VRAM
Very Tight
>95% VRAM
CPU Offload
VRAM+RAM split
Won't Fit
No memory

Model Compatibility

Estimates show the best-fit quantization for your VRAM. Expand any card for a full per-quant breakdown.

95 models · last updated Oct 2026

Top Picks for Your Hardware

How we calculate?

Weights = total parameters × llama.cpp bits-per-weight for each quant (Q4_K_M ≈ 4.9 bits, checked against real GGUF files). KV cache is computed from each model's own config.json — layers, KV heads, head size, sliding-window and linear-attention layers, MLA — for your context length. Plus 0.65–1.5 GB runtime overhead. On Macs, the GPU budget is ~75% of unified memory. For MoE models (Mixtral, Qwen 3 MoE, DeepSeek R1), all weights load into VRAM even though only a fraction are active per token. This is why Mixtral 8x7B shows ~27 GB, not ~13 GB — that's the accurate number for running it in Ollama or llama.cpp.

What is Offloading?

If a model is too big for your GPU VRAM, tools like Llama.cpp split it. Part runs fast on GPU, the rest runs slowly on System RAM (CPU). We mark these as "Doable (Slow)".

Privacy First

Like all id8 tools, this runs 100% in your browser. We don't collect your hardware data or usage stats.

GPU Recommendations by Model Size (2026)

Picking the right GPU for local LLM inference starts with knowing how much VRAM each model size class actually needs at a practical quantization level. The table below maps VRAM budgets to recommended consumer and prosumer GPUs, the model sizes that fit at Q4_K_M (the default quantization for most Ollama and llama.cpp setups), and realistic context window limits before KV cache overflow becomes a problem. Use this as a first-pass guide — your exact numbers will vary with runtime, driver version, and the specific model architecture.

VRAM Budget Recommended GPU(s) Model Sizes That Fit (Q4_K_M) Max Context (practical)
4–6 GB RTX 3060 6GB, RX 6600 1B–3B models (Phi-4-mini, Gemma 2B) 4K
8 GB RTX 3060 12GB, RTX 4060, M2/M3 8GB 3B–10B models (Llama 4 8B Q4, Gemma 4 4B) 8K
12 GB RTX 3080 12GB, RTX 4070 7B–13B models (Llama 3.1 8B, Qwen 3 8B) 16K
16 GB RTX 4080, M2 Pro 16GB, M3 Pro 18GB 13B–20B models, small MoE 32K
24 GB RTX 3090/4090, M4 Max 24GB 20B–34B models (Gemma 4 27B Q4, Qwen 3 32B Q3) 64K
48 GB RTX 6000 Ada, dual 3090 NVLink, M4 Max 48GB Up to 70B models (Llama 4 Maverick, DeepSeek R1 70B) 128K
80 GB+ A100 80GB, H100 80GB 70B FP16, large MoE active layers 128K+

Note on MoE architecture: Mixture-of-Experts models like Llama 4 Maverick, Mixtral 8x22B, and DeepSeek V3 have two parameter counts that matter: total parameters and active parameters per token. Only a subset of expert weights are activated per forward pass, which reduces compute — but all expert weights must still be resident in VRAM. This means a 400B-parameter MoE model still requires VRAM proportional to its total weight size, not just the active slice. Always check both counts when sizing your hardware for MoE deployments.

The table above reflects practical reality for single-GPU setups in 2026. Multi-GPU configurations unlock larger model sizes through tensor parallelism (splitting attention heads across GPUs) or pipeline parallelism (splitting model layers across GPUs), but these configurations require more complex setup and are typically reserved for server-grade deployments. For most individual developers, the single-GPU tiers above represent the practical ceiling. If your use case requires models above 70B parameters with full precision, you are likely looking at cloud inference APIs rather than local hardware — and the cost-benefit math for buying an A100 or H100 changes significantly compared to cloud API pricing.

One underappreciated category in this table is the Apple M-series unified memory tier. An M4 Max with 48GB unified memory offers VRAM bandwidth that rivals many NVIDIA cards, and the integration of CPU and GPU memory eliminates the PCIe bottleneck that makes CPU offload so slow on discrete GPU systems. For developers on macOS who need to run 30B–70B models without buying NVIDIA server hardware, a high-unified-memory Apple Silicon Mac is often the most cost-effective option available in the consumer market today.

VRAM Formula: How We Calculate Requirements

The VRAM estimate shown by this tool is not a single lookup — it is a three-component sum that combines weight memory, KV cache growth, and runtime overhead. Understanding each component helps you tune the estimate for your specific model, quantization, and context window, and explains why two models with similar parameter counts can have very different memory footprints in practice.

Total VRAM = Weight Memory + KV Cache + Runtime Overhead

Weight Memory = Total Parameters × Bits per Weight ÷ 8   (MoE: every expert is loaded)
  Q3_K_M  → ~3.95 bits  (~0.49 GB per 1B params)
  Q4_K_M  → ~4.88 bits  (~0.61 GB per 1B params)
  Q5_K_M  → ~5.70 bits  (~0.71 GB per 1B params)
  Q6_K    → ~6.57 bits  (~0.82 GB per 1B params)
  Q8_0    → ~8.50 bits  (~1.06 GB per 1B params)
  F16     → 16 bits     (~2.00 GB per 1B params)

KV Cache = 2 × full-attention layers × num_kv_heads × head_dim × context_length × bytes_per_element
  (sliding-window layers stop at their window; linear-attention / Mamba layers add none;
   MLA models — DeepSeek, Kimi — store one compressed latent per layer instead)
  Typical at   4K context: ~0.5–1.5 GB  (depends on architecture)
  Typical at  32K context: ~4–12 GB
  Typical at 128K context: ~16–50 GB

Runtime Overhead: ~0.5–1.5 GB (graph buffers, allocator, tokenizer state)

Worked example — Llama 4 Scout (MoE, 17B active of 109B total) at Q4_K_M, 32K context:

The active-parameter count sets speed, not memory: MoE models are fast for their size but need memory for every expert. Scout's attention design keeps the KV cache small (~192 KB per token at short context) because most of its layers only look back 8,192 tokens. Always sanity-check by running a short generation and watching GPU memory with nvidia-smi or Activity Monitor on macOS.

One important detail in the KV cache formula is the role of grouped-query attention (GQA). Models that use GQA — including most modern architectures like Llama 4, Qwen 3, and Gemma 4 — share KV heads across multiple query heads. This directly reduces KV cache size proportional to the ratio of query heads to KV heads. A model with 32 query heads and 8 KV heads has a KV cache that is 4× smaller than a multi-head attention model of the same layer depth. This is why newer models often have better long-context VRAM efficiency than older architectures of similar depth — not because their weights are smaller, but because their attention design is more memory-efficient. The tool accounts for this when it has GQA metadata available.

For CPU offload scenarios, the formula changes significantly. When layers are offloaded to system RAM, the weight memory for those layers moves off the GPU — but any layer that generates KV cache entries still contributes to GPU KV memory if you keep the KV cache on GPU. Mixed offload strategies that keep the attention layers on GPU and offload MLP layers to CPU can achieve better generation speed than full offload, but require careful layer-by-layer profiling. Tools like llama.cpp expose the -ngl flag to control how many layers stay on GPU, allowing you to tune the tradeoff between VRAM usage and generation speed for your specific hardware.

Quantization Formats: Q4_K_M vs Q5_K_M vs Q8_0 vs FP16

Quantization is the process of reducing the precision of model weights from the 32-bit or 16-bit floats used during training to lower-bit representations that are more memory-efficient at inference time. Choosing the right quantization format is one of the highest-leverage decisions in local LLM deployment: it directly controls how much VRAM you need, how fast tokens generate, and how much quality you sacrifice relative to the original model checkpoint.

Format Bits per Weight VRAM per 1B Params Quality vs FP16 Best For
Q4_K_M 4-bit (mixed) ~0.5 GB −2–4% Daily use, consumer GPUs
Q5_K_M 5-bit (mixed) ~0.62 GB −1–2% Better quality/VRAM balance
Q6_K 6-bit ~0.75 GB <1% loss Near-lossless, 24GB+ GPUs
Q8_0 8-bit ~1.0 GB <0.5% loss Research, 48GB+ GPUs
FP16 16-bit ~2.0 GB Baseline Full precision serving
GGUF IQ2_XS 2-bit IQ ~0.3 GB −5–10% Extremely constrained hardware

When to choose Q4_K_M (our default recommendation): For most local inference use cases on consumer hardware — gaming GPUs, MacBook Pro with M-series chips, or laptops with 8–16GB VRAM — Q4_K_M is the sweet spot. The 2–4% quality loss versus FP16 is imperceptible for code generation, summarization, and chat tasks. It lets you run a 7B model on an 8GB card with room for a reasonable context window, and it is the format most Ollama models default to. If you are unsure, start here.

When Q5_K_M or Q6_K is worth the extra VRAM: If you have 12–24GB of VRAM and are running tasks where output quality matters more than generation speed — structured data extraction, legal document analysis, translation with domain-specific terminology — the step up to Q5_K_M or Q6_K is worth approximately 25–50% more VRAM for a noticeable reduction in perplexity-measured quality loss. Q6_K is particularly well-regarded for instruction-following fidelity at smaller model sizes where every bit of precision counts.

When to avoid low-bit quantization (IQ2, IQ3, Q3): Sub-4-bit quantization formats like IQ2_XS can make previously-impossible model sizes technically loadable on constrained hardware, but the quality degradation is often severe enough to make the model unreliable for production tasks. Hallucination rates rise, instruction following degrades, and the model may fail to respect system prompts or output formats. These formats are best treated as a last resort for experimentation rather than deployment.

It is also worth noting that quantization quality is not uniform across model layers. The K-quant formats (Q4_K_M, Q5_K_M, Q6_K) use a mixed-precision strategy that applies higher precision to certain sensitive layers — typically the first and last transformer layers, and the attention projection matrices — while using lower precision for the bulk of the MLP weights. This is why K-quant formats outperform naive per-tensor quantization at the same bit width. When evaluating a model for a production task, always use the K-quant variant if it is available rather than the plain Q4_0 or Q5_0 variant.

For users who need to choose between a smaller model at higher precision and a larger model at lower precision with the same VRAM budget, the answer is almost always: larger model, lower precision. A 13B model at Q4_K_M will consistently outperform a 7B model at FP16 on reasoning and knowledge tasks, despite using similar VRAM. Model scale dominates quantization quality loss in most practical benchmarks below the FP16 precision floor. The exception is extremely low bit widths (2-bit), where quality loss becomes so severe that a smaller high-precision model may actually be more coherent.

Choosing an Inference Engine: Ollama vs llama.cpp vs LM Studio vs MLX vs vLLM

The inference engine you choose affects VRAM overhead, throughput, ease of setup, and which hardware backends are available to you. Most local inference stacks build on top of llama.cpp under the hood, but they differ significantly in API surface, memory management, and optimization for specific hardware. Here is a practical comparison of the most widely-used options in 2026.

Engine Best For Platform VRAM Overhead GPU Backend Notes
Ollama Beginners, CLI usage Linux/Mac/Win Low CUDA, Metal One-command setup, REST API
llama.cpp Max control, custom builds All Minimal CUDA, Metal, Vulkan, OpenBLAS Underlying engine for most tools
LM Studio GUI users, model hub Mac/Win Low CUDA, Metal Easy model browsing, built on llama.cpp
MLX Apple Silicon only macOS M-series Minimal Metal only Best throughput on M-series, native Swift/Python
vLLM Production serving, batching Linux (CUDA) Higher CUDA Continuous batching, PagedAttention, OpenAI-compatible API
TGI (HuggingFace) HuggingFace ecosystem Linux Medium CUDA Flash Attention, tensor parallelism for multi-GPU

Recommendation summary: For most individual developers running models locally, Ollama is the right starting point — it handles model download, GGUF conversion, and serving in a single command with sensible defaults. Mac users on M-series hardware should also evaluate MLX, which achieves significantly higher tokens-per-second throughput by using Apple's Metal framework natively rather than through a compatibility layer. For teams running a shared inference server or needing concurrent request handling with SLAs, vLLM is the production-grade choice on Linux with NVIDIA hardware. LM Studio is ideal for users who prefer a graphical interface and want to browse and compare models without touching a terminal. All of these engines ultimately rely on the same underlying model weights — the choice of engine does not change which models you can run, only how efficiently you can serve them and how much operational complexity you accept.

Common VRAM Mistakes (and How to Avoid Them)

1. Forgetting KV Cache Growth

KV cache grows linearly with context length, and this catches more developers off guard than any other VRAM issue. A model that fits comfortably in GPU memory at a 4K context window can run out of memory mid-generation at 32K because the key-value cache for attention layers expands to fill every additional token slot you allocate. Always check your intended context window, not just model size. If you need 64K or 128K context, factor that KV cache cost into your hardware selection before you commit to a model or a GPU. The worked formula example above shows how significant this gap can be: a model that needs 9GB at 4K might need 16GB+ at 32K.

2. Misunderstanding MoE Memory

Mixture-of-Experts models such as Mixtral 8x22B, Llama 4 Maverick, and DeepSeek V3 have large total parameter counts but smaller active counts per token. The compute cost scales with active parameters — which is why MoE models can match dense model quality at lower inference latency — but the VRAM requirement scales with total parameter count, because all expert weights must be loaded into memory before routing can select which subset to activate. A 400B MoE model with 17B active parameters per token still requires approximately 200GB VRAM to hold all weights at FP16. Advertising the active parameter count without the total count is a common source of confusion in benchmark writeups and product announcements.

3. Shared GPU Memory on Laptops

Laptop GPUs — including the RTX 4060 Mobile, RTX 4070 Mobile, and most gaming laptop variants — share a memory pool with system RAM. The 8GB or 16GB figure advertised in spec sheets is the total shared budget, not a dedicated GPU allocation. The operating system, display driver, and background applications all draw from the same pool, leaving the actual VRAM available for model weights at 6–7GB on an "8GB" system. When sizing for laptop deployment, subtract 1–2GB from the advertised figure and plan accordingly. Apple Silicon unified memory does not have this problem in the same way, because the entire unified memory pool is accessible to both CPU and GPU without contention.

4. Ignoring Runtime Overhead

Every inference runtime requires memory beyond the raw model weights. Allocators, CUDA graph buffers, tokenizer state, sampling tensors, and framework metadata together consume 0.5 to 1.5GB of VRAM that is completely separate from the model size or KV cache. A 7GB model on an 8GB GPU has essentially no headroom — it may load but fail mid-generation when the runtime tries to expand internal buffers. As a rule of thumb, always leave at least 1GB of unallocated VRAM as headroom for runtime overhead, and 2GB if you are using a GUI wrapper like LM Studio or Open WebUI on top of the inference backend.

5. Using FP32 by Mistake

Some tools and frameworks default to FP32 weight loading when a model is stored in safetensors format without an explicit quantization instruction. FP32 uses 4 bytes per parameter versus 2 bytes for FP16 or BF16 — exactly double the VRAM for identical model weights. A 7B model that should consume ~14GB at FP16 suddenly needs ~28GB, making it impossible to run on any consumer GPU. Always specify a quantization format or precision explicitly in your loading call or configuration file. In Ollama, set the quantization in the Modelfile. In llama.cpp, pass the correct GGUF file. In HuggingFace Transformers, use torch_dtype=torch.float16 or load_in_4bit=True explicitly rather than relying on defaults.

A practical diagnostic for this mistake: after loading a model, check the reported VRAM usage with nvidia-smi and compare it against the expected size for your quantization format from the table in the Quantization section above. If actual usage is roughly double the expected Q4_K_M estimate, the model likely loaded in FP32. The fix is to explicitly pass the desired dtype or to use a GGUF-format checkpoint which embeds quantization in the file format itself and cannot accidentally load in FP32. GGUF is the safest format for local inference precisely because it makes the precision explicit and non-overridable by framework defaults.

CanIRunLLM Technical FAQ and VRAM Planning Guide

What this tool does and why it exists

CanIRunLLM estimates whether a model will run on your hardware before you spend time downloading multi-gigabyte weights. It combines model size, quantization assumptions, baseline runtime overhead, and context window effects to estimate memory pressure in a way that matches how local runtimes behave in practice. The goal is not to produce a perfect lab-grade benchmark. The goal is to answer a practical question fast: will this model run at all, will it run only with RAM offload, or should you choose a smaller model and avoid a failed setup cycle.

Local LLM workflows often fail because people only look at parameter count. Parameter count matters, but the deployable memory footprint also depends on quantization format, runtime metadata, and KV cache growth as context grows. For example, two models with similar parameter counts can behave very differently if one is mixture-of-experts and activates fewer experts per token, or if one is tested at 4K context and the other at 128K context. This page translates those moving parts into a single compatibility view so developers can pick feasible configurations quickly.

Brief history of the problem

In early open model adoption, many setup guides used rough rules like "7B fits on 8 GB" and stopped there. That worked only for narrow cases. As local inference became mainstream, quantization formats multiplied, context windows expanded, and consumer hardware diversity increased. A single static table became unreliable. Developers began asking not just whether a model can load, but whether it can run at acceptable speed, whether long context is realistic, and whether offloading is worth it. The modern problem is capacity planning, not only model loading.

The mismatch between theoretical model size and observed runtime usage is the core source of confusion. Weight memory is only one component. Runtime allocators, tokenizer state, graph buffers, and KV cache consume additional memory. When those components are ignored, users overestimate what their GPU can handle and blame the runtime for crashes or poor throughput. Tooling like this emerged to encode practical sizing heuristics from real local inference behavior, especially for Ollama, llama.cpp-based stacks, and desktop GUIs such as LM Studio.

How to calculate this programmatically without the tool

You can build your own estimator in any language using a deterministic pipeline. First, collect model metadata: effective active parameters, architecture class, and supported quantization variants. Second, map quantization to bytes per parameter. Third, calculate weight memory as active_parameters multiplied by bytes_per_parameter. Fourth, add runtime overhead for graph buffers and allocator slack. Fifth, add KV cache memory based on context length, layer count, hidden size, and precision of cache tensors. Finally, compare the total against available VRAM and optionally against system RAM for offload scenarios.

A practical starter approximation for many Q4 local runs is to estimate around 0.5 to 0.6 GB per 1B active parameters, then add a base overhead buffer and a context-dependent cache term. If you are implementing this in code, keep the constants configurable instead of hardcoded. Runtime versions, drivers, and backend kernels shift real usage over time. You should also version your estimator assumptions and log estimator error versus observed usage, so your numbers improve after each deployment rather than remaining static guesses.

Engineering checklist for reliable local deployment

Separate compatibility from performance. A model that barely loads can still be unusable for interactive work. Add a throughput target to your planner, for example minimum tokens per second at a fixed prompt length. Include fallback recommendations in your script: reduce context, switch quantization, or downgrade model size. Detect unified-memory systems separately from discrete GPU systems because memory contention patterns differ. When offload is necessary, warn users that generation latency can rise sharply and that CPU memory bandwidth becomes the limiting factor.

Validate estimates with real runs. Your script should run a short warm-up inference and record peak memory and generation speed, then compare those values against the estimate. If divergence is large, update constants for that backend and hardware class. This approach turns compatibility checking into an iterative calibration process instead of a one-time spreadsheet. Over time you get a per-runtime and per-device memory model that predicts failures earlier and recommends viable alternatives before users waste time on failed downloads.

When to trust estimates and when to benchmark directly

Estimates are most reliable for common quantized checkpoints, moderate context windows, and single-user workloads. You should benchmark directly when you plan high context windows, concurrent sessions, speculative decoding, custom kernels, or production SLAs. Also benchmark when using new model families where effective active parameter behavior differs from old assumptions. The safest workflow is: estimate first for fast pruning, benchmark second for final validation. That two-step process is exactly the operational gap this tool is meant to close for local AI developers.

Sources & Methodology

VRAM estimates on this page are derived from empirical benchmarks and open community research. Key references:

Estimates are approximations. Actual VRAM usage varies by runtime version, driver, and quantization implementation. Always verify with nvidia-smi or Activity Monitor on Apple Silicon.

Related Tools

🎯 Can I Fine-Tune LLM? 🖥️ What GPU Do I Need? 🏆 LLM Leaderboard 💻 Best LLM for Coding 🧠 Best LLM for Reasoning 📐 Best LLM for Math 👁️ Best LLM for Vision 🇮🇳 Best LLM for Hindi

Related Guides

📖 Best GPU for Running LLM Locally 📖 How to Calculate VRAM Requirements 📖 LLM Quantization Q4 vs Q8 Explained 📖 Can You Run LLM With 8GB RAM? 📖 How to Run LLM with Ollama