Local AI Hardware
How to Run Gemma 4 Locally — E2B, E4B, 27B & 31B Guide (2026)
Last updated: 2026-05-07 — covers full Gemma 4 family released April 2026
Google released the Gemma 4 family in April 2026, and it's one of the most hardware-friendly multimodal model families available for local inference. The family spans from tiny 2.3B edge models that run on any laptop GPU all the way to a 31B dense model for workstations. All variants support vision (images in prompt), 128K context, and are Apache 2.0 licensed — meaning you can use them commercially without restrictions. This guide covers VRAM requirements, Ollama commands, and practical tips for every model in the family.
Developer's Note
Gemma 4 is notable for being genuinely multimodal at the edge scale. The E2B and E4B models accept image inputs and produce coherent descriptions even at 4-bit quantization, which is unusual for sub-5B models. If you're building a local vision pipeline on constrained hardware, start here.
The Gemma 4 Family at a Glance
The Gemma 4 lineup is split into two categories: edge models (E2B, E4B) optimized for constrained hardware, and full models (26B MoE, 31B dense) for workstations and servers. All share the same 128K context window and multimodal (image + text) input support.
| Model | Params | Type | Min VRAM (Q4) | Ollama Tag |
|---|---|---|---|---|
| Gemma 4 E2B | 2.3B | Edge, dense | ~1.8 GB | gemma4:e2b |
| Gemma 4 E4B | 4.5B | Edge, dense | ~3 GB | gemma4:e4b |
| Gemma 4 26B MoE | 26B (MoE) | MoE, ~4B active | ~8 GB | gemma4:26b |
| Gemma 4 31B | 31B | Dense | ~19 GB | gemma4:31b |
VRAM estimates at Q4_K_M quantization. Actual usage varies by context length and runtime overhead. The 26B MoE only loads ~4B active parameters per token during inference, which is why it fits in 8GB. Use Can I Run LLM for exact estimates on your hardware.
Running the Edge Models: E2B and E4B
The E2B and E4B models are the best starting point. They are specifically designed to run on mobile chips, laptop GPUs, and any system with 4GB+ VRAM. Both support image inputs, making them useful for local vision tasks that would otherwise require cloud APIs.
# Pull and run E2B (2.3B params — fits any 4GB+ GPU)
ollama pull gemma4:e2b
ollama run gemma4:e2b
# Pull and run E4B (4.5B params — recommended for quality/size balance on 4–8GB)
ollama pull gemma4:e4b
ollama run gemma4:e4b
To use the multimodal (vision) capability, pass an image path in the Ollama interactive session
with /path/to/image.jpg followed by your question, or use the API:
import base64, requests
with open("screenshot.png", "rb") as f:
img_b64 = base64.b64encode(f.read()).decode()
resp = requests.post(
"http://localhost:11434/v1/chat/completions",
json={
"model": "gemma4:e4b",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img_b64}"}},
{"type": "text", "text": "Describe what you see in this screenshot."}
]
}]
},
timeout=60,
)
print(resp.json()["choices"][0]["message"]["content"])
Running the 26B MoE Model
The 26B MoE model uses a Mixture-of-Experts architecture where only about 4B parameters are active per token during inference. This means it fits in 8GB VRAM at Q4_K_M while delivering quality closer to a 26B model than a 4B model on most tasks — particularly reasoning and longer outputs.
# Requires ~8GB VRAM at Q4_K_M
ollama pull gemma4:26b
ollama run gemma4:26b
If you are on a 16GB system (e.g. RTX 3080/4070 Ti), you can run the 26B MoE at Q8_0 for better output quality at roughly 14GB VRAM usage.
Running the 31B Dense Model
The 31B dense model is Gemma 4's flagship for local inference. It earned a Chatbot Arena ELO of 1452 and an MMLU-Pro score of 85.2% — competitive with much larger API models. However, it requires at least a 24GB VRAM GPU (RTX 3090, RTX 4090, or equivalent) at Q4_K_M to run fully in VRAM.
# Requires ~19GB VRAM at Q4_K_M — RTX 3090 / 4090 recommended
ollama pull gemma4:31b
ollama run gemma4:31b
# On 16GB VRAM, use CPU offload (slower but functional):
ollama run --num-gpu-layers 30 gemma4:31b
Which Gemma 4 Model Should You Pick?
| Your Setup | Recommended Model | Why |
|---|---|---|
| 4–6GB VRAM (laptop GPU, GTX 1060/1660) | E4B Q4_K_M | Best quality that fits without offload |
| 8GB VRAM (RTX 3060/4060) | 26B MoE Q4_K_M | MoE inference runs on ~4B active params — big quality jump |
| 16GB VRAM (RTX 3080/4070 Ti) | 26B MoE Q8_0 | Full quality MoE, or 31B with partial offload |
| 24GB+ VRAM (RTX 3090/4090) | 31B Q4_K_M | Flagship quality, Arena ELO 1452, fully in VRAM |
llama.cpp Alternative (Advanced Users)
If you need more control over quantization or GPU layer splits, build from the Hugging Face
GGUF files directly with llama.cpp. Gemma 4 GGUF files are available on Hugging Face under
google/gemma-4-*-it repositories (or community GGUF conversions).
# Example: run E4B with llama.cpp
./llama-cli \
-m gemma-4-e4b-it-Q4_K_M.gguf \
-n 512 \
--gpu-layers 35 \
-p "Describe the architecture of a transformer model."
Key Limitations
- Vision input token cost: Images are tokenized and counted against the 128K context. High-resolution images can consume thousands of tokens, so resize to 512–1024px for efficiency in production.
- MoE disk footprint: The 26B MoE model downloads all 26B parameters (~15GB on disk at Q4), even though only ~4B are active per inference. Plan disk usage accordingly.
- Quantization quality cliff: At Q2 quantization, Gemma 4 models exhibit noticeable quality degradation on reasoning tasks. Stick to Q4_K_M or higher for reliable outputs.
- 31B requires 24GB+ for full GPU inference: On 16GB VRAM, expect significant CPU offload and slower generation (2–5 tokens/sec instead of 20–30).
Check Your Hardware First
Before downloading multi-gigabyte model files, validate your setup with Can I Run LLM. Select the Gemma 4 variant you want and enter your GPU — it will calculate exact VRAM requirements at each quantization level.
Check if Gemma 4 runs on your GPU — free, no signup
Can I Run LLM — Free VRAM Calculator →