CanIRunLLM is a free VRAM calculator that checks if your GPU can run 95 self-hosted LLMs — Llama 4, Gemma 4, Kimi K2, DeepSeek R1. No upload, no signup.
Use the id8 CanIRunLLM calculator to check if your GPU VRAM is sufficient for local AI models in 2026 —
Llama 4 Maverick, Llama 4 Scout, Gemma 4 (12B, 26B, 31B), Kimi K2.7-Code, DeepSeek R1, Mistral Small 3.2,
Phi-4, Qwen 3, Nemotron 3 Ultra, MiniMax M2.7 and 50 more open-weight models.
Free in-browser VRAM calculator. No upload, no signup, no telemetry. Updated October 2026.
Your Hardware
12GB
4GB192GB
32GB
4GB256GB
*macOS lets the GPU use ~75% of unified memory: 13.5 GB of
18 GB. Raise it with sudo sysctl iogpu.wired_limit_mb (at your own risk).
* Used for "Offloading" when model doesn't fit in VRAM. Slower than GPU only.
4096Tokens
4K128K+
* Longer context requires more VRAM for the KV cache.
Runs Great
<60% VRAM
Runs Well
60–80% VRAM
Tight Fit
80–95% VRAM
Very Tight
>95% VRAM
CPU Offload
VRAM+RAM split
Won't Fit
No memory
Model Compatibility
Estimates show the best-fit quantization for your VRAM. Expand any card for a full per-quant breakdown.
95 models · last updated Oct 2026
Math: all params × GGUF bits/weight (Q4_K_M ≈ 4.9)
+ exact KV cache (config.json) + runtime overhead
Top Picks for Your Hardware
How we calculate?
Weights = total parameters × llama.cpp bits-per-weight for each quant (Q4_K_M ≈ 4.9 bits,
checked against real GGUF files). KV cache is computed from each model's own config.json —
layers, KV heads, head size, sliding-window and linear-attention layers, MLA — for your
context length. Plus 0.65–1.5 GB runtime overhead. On Macs, the GPU budget is ~75% of unified memory. For MoE models (Mixtral, Qwen 3 MoE, DeepSeek R1), all weights
load into VRAM even though only a fraction are active per token. This is why Mixtral 8x7B
shows ~27 GB, not ~13 GB — that's the accurate number for running it in Ollama or llama.cpp.
What is Offloading?
If a model is too big for your GPU VRAM, tools like Llama.cpp split it. Part runs fast on GPU,
the rest runs slowly on System RAM (CPU). We mark these as "Doable (Slow)".
Privacy First
Like all id8 tools, this runs 100% in your browser. We don't collect your hardware data or usage
stats.
GPU Recommendations by Model Size (2026)
Picking the right GPU for local LLM inference starts with knowing how much VRAM each model size class
actually needs at a practical quantization level. The table below maps VRAM budgets to recommended
consumer and prosumer GPUs, the model sizes that fit at Q4_K_M (the default quantization for most
Ollama and llama.cpp setups), and realistic context window limits before KV cache overflow becomes a
problem. Use this as a first-pass guide — your exact numbers will vary with runtime, driver version,
and the specific model architecture.
VRAM Budget
Recommended GPU(s)
Model Sizes That Fit (Q4_K_M)
Max Context (practical)
4–6 GB
RTX 3060 6GB, RX 6600
1B–3B models (Phi-4-mini, Gemma 2B)
4K
8 GB
RTX 3060 12GB, RTX 4060, M2/M3 8GB
3B–10B models (Llama 4 8B Q4, Gemma 4 4B)
8K
12 GB
RTX 3080 12GB, RTX 4070
7B–13B models (Llama 3.1 8B, Qwen 3 8B)
16K
16 GB
RTX 4080, M2 Pro 16GB, M3 Pro 18GB
13B–20B models, small MoE
32K
24 GB
RTX 3090/4090, M4 Max 24GB
20B–34B models (Gemma 4 27B Q4, Qwen 3 32B Q3)
64K
48 GB
RTX 6000 Ada, dual 3090 NVLink, M4 Max 48GB
Up to 70B models (Llama 4 Maverick, DeepSeek R1 70B)
128K
80 GB+
A100 80GB, H100 80GB
70B FP16, large MoE active layers
128K+
Note on MoE architecture: Mixture-of-Experts models like
Llama 4 Maverick, Mixtral 8x22B, and DeepSeek V3 have two parameter counts that matter: total
parameters and active parameters per token. Only a subset of expert weights are activated per
forward pass, which reduces compute — but all expert weights must still be resident in VRAM. This
means a 400B-parameter MoE model still requires VRAM proportional to its total weight size, not
just the active slice. Always check both counts when sizing your hardware for MoE deployments.
The table above reflects practical reality for single-GPU setups in 2026. Multi-GPU configurations
unlock larger model sizes through tensor parallelism (splitting attention heads across GPUs) or
pipeline parallelism (splitting model layers across GPUs), but these configurations require more
complex setup and are typically reserved for server-grade deployments. For most individual
developers, the single-GPU tiers above represent the practical ceiling. If your use case requires
models above 70B parameters with full precision, you are likely looking at cloud inference APIs
rather than local hardware — and the cost-benefit math for buying an A100 or H100 changes
significantly compared to cloud API pricing.
One underappreciated category in this table is the Apple M-series unified memory tier. An M4 Max
with 48GB unified memory offers VRAM bandwidth that rivals many NVIDIA cards, and the integration
of CPU and GPU memory eliminates the PCIe bottleneck that makes CPU offload so slow on discrete
GPU systems. For developers on macOS who need to run 30B–70B models without buying NVIDIA server
hardware, a high-unified-memory Apple Silicon Mac is often the most cost-effective option
available in the consumer market today.
VRAM Formula: How We Calculate Requirements
The VRAM estimate shown by this tool is not a single lookup — it is a three-component sum that
combines weight memory, KV cache growth, and runtime overhead. Understanding each component helps
you tune the estimate for your specific model, quantization, and context window, and explains why
two models with similar parameter counts can have very different memory footprints in practice.
Total VRAM = Weight Memory + KV Cache + Runtime Overhead
Weight Memory = Total Parameters × Bits per Weight ÷ 8 (MoE: every expert is loaded)
Q3_K_M → ~3.95 bits (~0.49 GB per 1B params)
Q4_K_M → ~4.88 bits (~0.61 GB per 1B params)
Q5_K_M → ~5.70 bits (~0.71 GB per 1B params)
Q6_K → ~6.57 bits (~0.82 GB per 1B params)
Q8_0 → ~8.50 bits (~1.06 GB per 1B params)
F16 → 16 bits (~2.00 GB per 1B params)
KV Cache = 2 × full-attention layers × num_kv_heads × head_dim × context_length × bytes_per_element
(sliding-window layers stop at their window; linear-attention / Mamba layers add none;
MLA models — DeepSeek, Kimi — store one compressed latent per layer instead)
Typical at 4K context: ~0.5–1.5 GB (depends on architecture)
Typical at 32K context: ~4–12 GB
Typical at 128K context: ~16–50 GB
Runtime Overhead: ~0.5–1.5 GB (graph buffers, allocator, tokenizer state)
Worked example — Llama 4 Scout (MoE, 17B active of 109B total) at Q4_K_M, 32K context:
Weight memory: all 109B params × 4.88 bits ÷ 8 = ~61.9 GB — every expert is loaded, even though only 17B are used per token
Total: approximately 66.0 GB — an 80 GB GPU, two 48 GB cards, or a 96 GB+ Mac
The active-parameter count sets speed, not memory: MoE models are fast for their size but need memory
for every expert. Scout's attention design keeps the KV cache small (~192 KB per token at
short context) because most of its layers only look back 8,192 tokens.
Always sanity-check by running a short generation and watching GPU memory with nvidia-smi
or Activity Monitor on macOS.
One important detail in the KV cache formula is the role of grouped-query attention (GQA). Models
that use GQA — including most modern architectures like Llama 4, Qwen 3, and Gemma 4 — share KV
heads across multiple query heads. This directly reduces KV cache size proportional to the ratio
of query heads to KV heads. A model with 32 query heads and 8 KV heads has a KV cache that is
4× smaller than a multi-head attention model of the same layer depth. This is why newer models
often have better long-context VRAM efficiency than older architectures of similar depth — not
because their weights are smaller, but because their attention design is more memory-efficient.
The tool accounts for this when it has GQA metadata available.
For CPU offload scenarios, the formula changes significantly. When layers are offloaded to system
RAM, the weight memory for those layers moves off the GPU — but any layer that generates KV cache
entries still contributes to GPU KV memory if you keep the KV cache on GPU. Mixed offload
strategies that keep the attention layers on GPU and offload MLP layers to CPU can achieve better
generation speed than full offload, but require careful layer-by-layer profiling. Tools like
llama.cpp expose the -ngl
flag to control how many layers stay on GPU, allowing you to tune the tradeoff between VRAM usage
and generation speed for your specific hardware.
Quantization Formats: Q4_K_M vs Q5_K_M vs Q8_0 vs FP16
Quantization is the process of reducing the precision of model weights from the 32-bit or 16-bit
floats used during training to lower-bit representations that are more memory-efficient at inference
time. Choosing the right quantization format is one of the highest-leverage decisions in local LLM
deployment: it directly controls how much VRAM you need, how fast tokens generate, and how much
quality you sacrifice relative to the original model checkpoint.
Format
Bits per Weight
VRAM per 1B Params
Quality vs FP16
Best For
Q4_K_M
4-bit (mixed)
~0.5 GB
−2–4%
Daily use, consumer GPUs
Q5_K_M
5-bit (mixed)
~0.62 GB
−1–2%
Better quality/VRAM balance
Q6_K
6-bit
~0.75 GB
<1% loss
Near-lossless, 24GB+ GPUs
Q8_0
8-bit
~1.0 GB
<0.5% loss
Research, 48GB+ GPUs
FP16
16-bit
~2.0 GB
Baseline
Full precision serving
GGUF IQ2_XS
2-bit IQ
~0.3 GB
−5–10%
Extremely constrained hardware
When to choose Q4_K_M (our default recommendation): For
most local inference use cases on consumer hardware — gaming GPUs, MacBook Pro with M-series chips,
or laptops with 8–16GB VRAM — Q4_K_M is the sweet spot. The 2–4% quality loss versus FP16 is
imperceptible for code generation, summarization, and chat tasks. It lets you run a 7B model on an
8GB card with room for a reasonable context window, and it is the format most Ollama models default
to. If you are unsure, start here.
When Q5_K_M or Q6_K is worth the extra VRAM: If you have
12–24GB of VRAM and are running tasks where output quality matters more than generation speed —
structured data extraction, legal document analysis, translation with domain-specific terminology —
the step up to Q5_K_M or Q6_K is worth approximately 25–50% more VRAM for a noticeable reduction
in perplexity-measured quality loss. Q6_K is particularly well-regarded for instruction-following
fidelity at smaller model sizes where every bit of precision counts.
When to avoid low-bit quantization (IQ2, IQ3, Q3):
Sub-4-bit quantization formats like IQ2_XS can make previously-impossible model sizes technically
loadable on constrained hardware, but the quality degradation is often severe enough to make the
model unreliable for production tasks. Hallucination rates rise, instruction following degrades,
and the model may fail to respect system prompts or output formats. These formats are best treated
as a last resort for experimentation rather than deployment.
It is also worth noting that quantization quality is not uniform across model layers. The
K-quant formats (Q4_K_M, Q5_K_M, Q6_K) use a mixed-precision strategy that applies higher
precision to certain sensitive layers — typically the first and last transformer layers, and the
attention projection matrices — while using lower precision for the bulk of the MLP weights. This
is why K-quant formats outperform naive per-tensor quantization at the same bit width. When
evaluating a model for a production task, always use the K-quant variant if it is available
rather than the plain Q4_0 or Q5_0 variant.
For users who need to choose between a smaller model at higher precision and a larger model at
lower precision with the same VRAM budget, the answer is almost always: larger model, lower
precision. A 13B model at Q4_K_M will consistently outperform a 7B model at FP16 on reasoning
and knowledge tasks, despite using similar VRAM. Model scale dominates quantization quality loss
in most practical benchmarks below the FP16 precision floor. The exception is extremely low bit
widths (2-bit), where quality loss becomes so severe that a smaller high-precision model may
actually be more coherent.
Choosing an Inference Engine: Ollama vs llama.cpp vs LM Studio vs MLX vs vLLM
The inference engine you choose affects VRAM overhead, throughput, ease of setup, and which
hardware backends are available to you. Most local inference stacks build on top of
llama.cpp under the hood, but they differ significantly
in API surface, memory management, and optimization for specific hardware. Here is a practical
comparison of the most widely-used options in 2026.
Engine
Best For
Platform
VRAM Overhead
GPU Backend
Notes
Ollama
Beginners, CLI usage
Linux/Mac/Win
Low
CUDA, Metal
One-command setup, REST API
llama.cpp
Max control, custom builds
All
Minimal
CUDA, Metal, Vulkan, OpenBLAS
Underlying engine for most tools
LM Studio
GUI users, model hub
Mac/Win
Low
CUDA, Metal
Easy model browsing, built on llama.cpp
MLX
Apple Silicon only
macOS M-series
Minimal
Metal only
Best throughput on M-series, native Swift/Python
vLLM
Production serving, batching
Linux (CUDA)
Higher
CUDA
Continuous batching, PagedAttention, OpenAI-compatible API
TGI (HuggingFace)
HuggingFace ecosystem
Linux
Medium
CUDA
Flash Attention, tensor parallelism for multi-GPU
Recommendation summary: For most individual developers
running models locally, Ollama is the right starting point — it handles model
download, GGUF conversion, and serving in a single command with sensible defaults. Mac users on
M-series hardware should also evaluate MLX, which achieves significantly higher
tokens-per-second throughput by using Apple's Metal framework natively rather than through a
compatibility layer. For teams running a shared inference server or needing concurrent request
handling with SLAs, vLLM is the production-grade choice on Linux with NVIDIA
hardware. LM Studio is ideal for users who prefer a graphical interface and want to browse and
compare models without touching a terminal. All of these engines ultimately rely on the same
underlying model weights — the choice of engine does not change which models you can run, only
how efficiently you can serve them and how much operational complexity you accept.
Common VRAM Mistakes (and How to Avoid Them)
1. Forgetting KV Cache Growth
KV cache grows linearly with context length, and this catches more developers off guard than any
other VRAM issue. A model that fits comfortably in GPU memory at a 4K context window can run out
of memory mid-generation at 32K because the key-value cache for attention layers expands to fill
every additional token slot you allocate. Always check your intended context window, not just
model size. If you need 64K or 128K context, factor that KV cache cost into your hardware
selection before you commit to a model or a GPU. The worked formula example above shows how
significant this gap can be: a model that needs 9GB at 4K might need 16GB+ at 32K.
2. Misunderstanding MoE Memory
Mixture-of-Experts models such as Mixtral 8x22B, Llama 4 Maverick, and DeepSeek V3 have large
total parameter counts but smaller active counts per token. The compute cost scales with active
parameters — which is why MoE models can match dense model quality at lower inference latency —
but the VRAM requirement scales with total parameter count, because all expert weights must be
loaded into memory before routing can select which subset to activate. A 400B MoE model with 17B
active parameters per token still requires approximately 200GB VRAM to hold all weights at FP16.
Advertising the active parameter count without the total count is a common source of confusion
in benchmark writeups and product announcements.
3. Shared GPU Memory on Laptops
Laptop GPUs — including the RTX 4060 Mobile, RTX 4070 Mobile, and most gaming laptop variants —
share a memory pool with system RAM. The 8GB or 16GB figure advertised in spec sheets is the
total shared budget, not a dedicated GPU allocation. The operating system, display driver, and
background applications all draw from the same pool, leaving the actual VRAM available for
model weights at 6–7GB on an "8GB" system. When sizing for laptop deployment, subtract 1–2GB
from the advertised figure and plan accordingly. Apple Silicon unified memory does not have this
problem in the same way, because the entire unified memory pool is accessible to both CPU and
GPU without contention.
4. Ignoring Runtime Overhead
Every inference runtime requires memory beyond the raw model weights. Allocators, CUDA graph
buffers, tokenizer state, sampling tensors, and framework metadata together consume 0.5 to 1.5GB
of VRAM that is completely separate from the model size or KV cache. A 7GB model on an 8GB GPU
has essentially no headroom — it may load but fail mid-generation when the runtime tries to
expand internal buffers. As a rule of thumb, always leave at least 1GB of unallocated VRAM as
headroom for runtime overhead, and 2GB if you are using a GUI wrapper like LM Studio or Open
WebUI on top of the inference backend.
5. Using FP32 by Mistake
Some tools and frameworks default to FP32 weight loading when a model is stored in safetensors
format without an explicit quantization instruction. FP32 uses 4 bytes per parameter versus 2
bytes for FP16 or BF16 — exactly double the VRAM for identical model weights. A 7B model that
should consume ~14GB at FP16 suddenly needs ~28GB, making it impossible to run on any consumer
GPU. Always specify a quantization format or precision explicitly in your loading call or
configuration file. In Ollama, set the quantization in the Modelfile. In llama.cpp, pass the
correct GGUF file. In HuggingFace Transformers, use
torch_dtype=torch.float16 or
load_in_4bit=True explicitly rather than relying
on defaults.
A practical diagnostic for this mistake: after loading a model, check the reported VRAM usage
with nvidia-smi and compare it against the
expected size for your quantization format from the table in the Quantization section above. If
actual usage is roughly double the expected Q4_K_M estimate, the model likely loaded in FP32. The
fix is to explicitly pass the desired dtype or to use a GGUF-format checkpoint which embeds
quantization in the file format itself and cannot accidentally load in FP32. GGUF is the safest
format for local inference precisely because it makes the precision explicit and non-overridable
by framework defaults.
CanIRunLLM Technical FAQ and VRAM Planning Guide
What this tool does and why it exists
CanIRunLLM estimates whether a model will run on your hardware before you spend time downloading
multi-gigabyte weights. It combines model size, quantization assumptions, baseline runtime overhead, and
context window effects to estimate memory pressure in a way that matches how local runtimes behave in
practice. The goal is not to produce a perfect lab-grade benchmark. The goal is to answer a practical
question fast: will this model run at all, will it run only with RAM offload, or should you choose a
smaller model and avoid a failed setup cycle.
Local LLM workflows often fail because people only look at parameter count. Parameter count matters, but
the deployable memory footprint also depends on quantization format, runtime metadata, and KV cache
growth as context grows. For example, two models with similar parameter counts can behave very
differently if one is mixture-of-experts and activates fewer experts per token, or if one is tested at
4K context and the other at 128K context. This page translates those moving parts into a single
compatibility view so developers can pick feasible configurations quickly.
Brief history of the problem
In early open model adoption, many setup guides used rough rules like "7B fits on 8 GB" and stopped
there. That worked only for narrow cases. As local inference became mainstream, quantization formats
multiplied, context windows expanded, and consumer hardware diversity increased. A single static table
became unreliable. Developers began asking not just whether a model can load, but whether it can run at
acceptable speed, whether long context is realistic, and whether offloading is worth it. The modern
problem is capacity planning, not only model loading.
The mismatch between theoretical model size and observed runtime usage is the core source of confusion.
Weight memory is only one component. Runtime allocators, tokenizer state, graph buffers, and KV cache
consume additional memory. When those components are ignored, users overestimate what their GPU can
handle and blame the runtime for crashes or poor throughput. Tooling like this emerged to encode
practical sizing heuristics from real local inference behavior, especially for Ollama, llama.cpp-based
stacks, and desktop GUIs such as LM Studio.
How to calculate this programmatically without the
tool
You can build your own estimator in any language using a deterministic pipeline. First, collect model
metadata: effective active parameters, architecture class, and supported quantization variants. Second,
map quantization to bytes per parameter. Third, calculate weight memory as active_parameters multiplied
by bytes_per_parameter. Fourth, add runtime overhead for graph buffers and allocator slack. Fifth, add
KV cache memory based on context length, layer count, hidden size, and precision of cache tensors.
Finally, compare the total against available VRAM and optionally against system RAM for offload
scenarios.
A practical starter approximation for many Q4 local runs is to estimate around 0.5 to 0.6 GB per 1B
active parameters, then add a base overhead buffer and a context-dependent cache term. If you are
implementing this in code, keep the constants configurable instead of hardcoded. Runtime versions,
drivers, and backend kernels shift real usage over time. You should also version your estimator
assumptions and log estimator error versus observed usage, so your numbers improve after each deployment
rather than remaining static guesses.
Engineering checklist for reliable local deployment
Separate compatibility from performance. A model that barely loads can still be unusable for interactive
work. Add a throughput target to your planner, for example minimum tokens per second at a fixed prompt
length. Include fallback recommendations in your script: reduce context, switch quantization, or
downgrade model size. Detect unified-memory systems separately from discrete GPU systems because memory
contention patterns differ. When offload is necessary, warn users that generation latency can rise
sharply and that CPU memory bandwidth becomes the limiting factor.
Validate estimates with real runs. Your script should run a short warm-up inference and record peak
memory and generation speed, then compare those values against the estimate. If divergence is large,
update constants for that backend and hardware class. This approach turns compatibility checking into an
iterative calibration process instead of a one-time spreadsheet. Over time you get a per-runtime and
per-device memory model that predicts failures earlier and recommends viable alternatives before users
waste time on failed downloads.
When to trust estimates and when to benchmark
directly
Estimates are most reliable for common quantized checkpoints, moderate context windows, and single-user
workloads. You should benchmark directly when you plan high context windows, concurrent sessions,
speculative decoding, custom kernels, or production SLAs. Also benchmark when using new model families
where effective active parameter behavior differs from old assumptions. The safest workflow is: estimate
first for fast pruning, benchmark second for final validation. That two-step process is exactly the
operational gap this tool is meant to close for local AI developers.
Sources & Methodology
VRAM estimates on this page are derived from empirical benchmarks and open community research. Key references:
llama.cpp — runtime memory allocation behaviour and CUDA/Metal overhead constants
Published GGUF files — real Q3–Q8 file sizes used to check the bits-per-weight figures (±5%)
Estimates are approximations. Actual VRAM usage varies by runtime version, driver, and quantization implementation. Always verify with nvidia-smi or Activity Monitor on Apple Silicon.