How-To

How to Run Llama 4 Locally: Hardware Needed, Ollama Commands and Settings

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

Llama 4 is Meta's open-weight family, and both models in it are large mixture-of-experts models. That makes the hardware question the first one to answer. This guide gives the memory figures, then the steps to run it with Ollama.

Every quantization and context length, for your exact GPU or Mac.

Check Llama 4 on your hardware →

The two models

Llama 4 ScoutLlama 4 Maverick
Parameters109B total, 17B active400B total, 17B active
Context window10M1M
Hugging Face repositorymeta-llama/Llama-4-Scout-17B-16E-Instructmeta-llama/Llama-4-Maverick-17B-128E-Instruct
Ollama tagllama4:scoutllama4:maverick
VRAM at Q4_K_M, 8K context64.9 GB230.2 GB

Both are mixture-of-experts models. Memory follows the total parameter count, since every expert has to be loaded; the active count affects speed. See MoE Models and VRAM.

Memory by quantization: Llama 4 Scout

QuantizationVRAM at 4K contextVRAM at 8K contextVRAM at 32K context
Q3_K_M52.4 GB53.1 GB54.2 GB
Q4_K_M64.2 GB64.9 GB66.0 GB
Q5_K_M74.6 GB75.3 GB76.5 GB
Q6_K85.6 GB86.4 GB87.5 GB
Q8_0110.1 GB110.9 GB112.0 GB

The context column matters. Llama 4 supports a very long context, but each token of context costs memory. Start with 8K and raise it only if you need to.

What hardware that implies

  • A single 24 GB or 32 GB GPU cannot hold Scout at Q4. It will run with layers offloaded to system RAM, at a few tokens per second.
  • Several GPUs are needed on a PC: add up their VRAM and compare it with the table above.
  • A Mac with 96 GB or 128 GB of unified memory holds Scout comfortably, since about three quarters of the memory is usable for the model.
  • Maverick is in workstation and server territory.

Check your own machine in Can I Run LLM?, or ask What GPU Do I Need? for the cheapest setup that fits.

Run it with Ollama

  1. Install Ollama from ollama.com. It is available for macOS, Windows and Linux.
  2. Pull and run the model:

ollama run llama4:scout

The first run downloads the weights, which are tens of gigabytes. Make sure you have the disk space.

  1. Chat in the terminal, or use the local API that Ollama serves on port 11434. Many chat front ends can connect to it.
  1. Set the context length. Ollama uses a modest default context. To change it inside a session:

/set parameter num_ctx 8192

Raise it only as far as your memory allows.

Run it with LM Studio

LM Studio is a desktop app with a model browser. Search for Llama 4 Scout, pick a quantization that fits your memory (the app shows an estimate), download and load it. On a Mac, LM Studio can also use MLX builds of the model.

A comparison of the two tools is in Ollama vs LM Studio.

Settings that matter

  • Quantization. Q4_K_M is the usual balance. Go lower only if you must; quality drops faster below Q4.
  • Context length. The biggest lever on memory after quantization.
  • GPU layers. If the model does not fit, the tool places some layers on the CPU. Fewer GPU layers means slower output.
  • Images. Llama 4 accepts images as well as text. Image input uses additional memory.

If it is too big for your machine

Most people do not have the hardware for Llama 4, and that is fine. Strong smaller models exist at every memory size:

Licence

Llama 4 is released under Meta's own community licence, not a standard open-source licence. It permits wide use with conditions. Read it on the model's Hugging Face page before building a product on it.

Frequently asked questions

How much VRAM does Llama 4 Scout need?

At Q4_K_M with an 8K context, Llama 4 Scout needs about 64.9 GB. It has 109B total, 17B active parameters.

Can I run Llama 4 on an RTX 4090?

Not fully in VRAM. Llama 4 Scout at Q4 needs far more than 24 GB. It can run with most layers in system RAM, slowly. A Mac with 96 GB or more of unified memory, or a multi-GPU machine, is the practical route.

What is the Ollama command for Llama 4?

For Scout: ollama run llama4:scout. For Maverick: ollama run llama4:maverick.

Is Llama 4 free to use?

The weights are free to download under Meta's Llama 4 Community License, which has its own conditions. Read the licence on the model page before commercial use.

What should I run if Llama 4 does not fit?

A smaller model. See the tables in our guide to the best local LLM for each VRAM size.

Related