Local AI Hardware

How to Run Gemma 4 Locally — E2B, E4B, 27B & 31B Guide (2026)

Last updated: 2026-05-07 — covers full Gemma 4 family released April 2026

Google released the Gemma 4 family in April 2026, and it's one of the most hardware-friendly multimodal model families available for local inference. The family spans from tiny 2.3B edge models that run on any laptop GPU all the way to a 31B dense model for workstations. All variants support vision (images in prompt), 128K context, and are Apache 2.0 licensed — meaning you can use them commercially without restrictions. This guide covers VRAM requirements, Ollama commands, and practical tips for every model in the family.

Developer's Note

Gemma 4 is notable for being genuinely multimodal at the edge scale. The E2B and E4B models accept image inputs and produce coherent descriptions even at 4-bit quantization, which is unusual for sub-5B models. If you're building a local vision pipeline on constrained hardware, start here.

The Gemma 4 Family at a Glance

The Gemma 4 lineup is split into two categories: edge models (E2B, E4B) optimized for constrained hardware, and full models (26B MoE, 31B dense) for workstations and servers. All share the same 128K context window and multimodal (image + text) input support.

Model Params Type Min VRAM (Q4) Ollama Tag
Gemma 4 E2B2.3BEdge, dense~1.8 GBgemma4:e2b
Gemma 4 E4B4.5BEdge, dense~3 GBgemma4:e4b
Gemma 4 26B MoE26B (MoE)MoE, ~4B active~8 GBgemma4:26b
Gemma 4 31B31BDense~19 GBgemma4:31b

VRAM estimates at Q4_K_M quantization. Actual usage varies by context length and runtime overhead. The 26B MoE only loads ~4B active parameters per token during inference, which is why it fits in 8GB. Use Can I Run LLM for exact estimates on your hardware.

Running the Edge Models: E2B and E4B

The E2B and E4B models are the best starting point. They are specifically designed to run on mobile chips, laptop GPUs, and any system with 4GB+ VRAM. Both support image inputs, making them useful for local vision tasks that would otherwise require cloud APIs.

# Pull and run E2B (2.3B params — fits any 4GB+ GPU)
ollama pull gemma4:e2b
ollama run gemma4:e2b

# Pull and run E4B (4.5B params — recommended for quality/size balance on 4–8GB)
ollama pull gemma4:e4b
ollama run gemma4:e4b

To use the multimodal (vision) capability, pass an image path in the Ollama interactive session with /path/to/image.jpg followed by your question, or use the API:

import base64, requests

with open("screenshot.png", "rb") as f:
    img_b64 = base64.b64encode(f.read()).decode()

resp = requests.post(
    "http://localhost:11434/v1/chat/completions",
    json={
        "model": "gemma4:e4b",
        "messages": [{
            "role": "user",
            "content": [
                {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img_b64}"}},
                {"type": "text", "text": "Describe what you see in this screenshot."}
            ]
        }]
    },
    timeout=60,
)
print(resp.json()["choices"][0]["message"]["content"])

Running the 26B MoE Model

The 26B MoE model uses a Mixture-of-Experts architecture where only about 4B parameters are active per token during inference. This means it fits in 8GB VRAM at Q4_K_M while delivering quality closer to a 26B model than a 4B model on most tasks — particularly reasoning and longer outputs.

# Requires ~8GB VRAM at Q4_K_M
ollama pull gemma4:26b
ollama run gemma4:26b

If you are on a 16GB system (e.g. RTX 3080/4070 Ti), you can run the 26B MoE at Q8_0 for better output quality at roughly 14GB VRAM usage.

Running the 31B Dense Model

The 31B dense model is Gemma 4's flagship for local inference. It earned a Chatbot Arena ELO of 1452 and an MMLU-Pro score of 85.2% — competitive with much larger API models. However, it requires at least a 24GB VRAM GPU (RTX 3090, RTX 4090, or equivalent) at Q4_K_M to run fully in VRAM.

# Requires ~19GB VRAM at Q4_K_M — RTX 3090 / 4090 recommended
ollama pull gemma4:31b
ollama run gemma4:31b

# On 16GB VRAM, use CPU offload (slower but functional):
ollama run --num-gpu-layers 30 gemma4:31b

Which Gemma 4 Model Should You Pick?

Your Setup Recommended Model Why
4–6GB VRAM (laptop GPU, GTX 1060/1660)E4B Q4_K_MBest quality that fits without offload
8GB VRAM (RTX 3060/4060)26B MoE Q4_K_MMoE inference runs on ~4B active params — big quality jump
16GB VRAM (RTX 3080/4070 Ti)26B MoE Q8_0Full quality MoE, or 31B with partial offload
24GB+ VRAM (RTX 3090/4090)31B Q4_K_MFlagship quality, Arena ELO 1452, fully in VRAM

llama.cpp Alternative (Advanced Users)

If you need more control over quantization or GPU layer splits, build from the Hugging Face GGUF files directly with llama.cpp. Gemma 4 GGUF files are available on Hugging Face under google/gemma-4-*-it repositories (or community GGUF conversions).

# Example: run E4B with llama.cpp
./llama-cli \
  -m gemma-4-e4b-it-Q4_K_M.gguf \
  -n 512 \
  --gpu-layers 35 \
  -p "Describe the architecture of a transformer model."

Key Limitations

  • Vision input token cost: Images are tokenized and counted against the 128K context. High-resolution images can consume thousands of tokens, so resize to 512–1024px for efficiency in production.
  • MoE disk footprint: The 26B MoE model downloads all 26B parameters (~15GB on disk at Q4), even though only ~4B are active per inference. Plan disk usage accordingly.
  • Quantization quality cliff: At Q2 quantization, Gemma 4 models exhibit noticeable quality degradation on reasoning tasks. Stick to Q4_K_M or higher for reliable outputs.
  • 31B requires 24GB+ for full GPU inference: On 16GB VRAM, expect significant CPU offload and slower generation (2–5 tokens/sec instead of 20–30).

Check Your Hardware First

Before downloading multi-gigabyte model files, validate your setup with Can I Run LLM. Select the Gemma 4 variant you want and enter your GPU — it will calculate exact VRAM requirements at each quantization level.

Check if Gemma 4 runs on your GPU — free, no signup

Can I Run LLM — Free VRAM Calculator →

Related Guides