How-To

How to Run Qwen 3 Locally: Which Size Fits Your GPU, and How to Start

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

Qwen 3, from Alibaba, is one of the easiest model families to run locally because it comes in many sizes, from one that fits on a phone to one that needs a server. This guide helps you pick the size for your hardware and get it running.

Every Qwen 3 size, every quantization, your exact hardware.

Check Qwen 3 on your GPU →

The family at a glance

ModelParametersVRAM at Q4_K_M, 8K contextOllama tag
Qwen 3 0.6B0.6B1.9 GBqwen3:0.6b
Qwen 3 1.7B1.7B2.5 GBqwen3:1.7b
Qwen 3 4B4B4.2 GBqwen3:4b
Qwen 3 8B8B6.5 GBqwen3:8b
Qwen 3 14B14.77B10.6 GBqwen3:14b
Qwen 3 30B-A3B30B total, 3B active18.8 GBqwen3:30b-a3b
Qwen 3 32B32B21.2 GBqwen3:32b
Qwen 3 235B-A22B235B total, 22.1B active136.5 GBqwen3:235b-a22b

Figures come from each model's published architecture through the same engine as Can I Run LLM?.

Pick a size for your hardware

Your memoryStart with
4 to 6 GB VRAM, or CPU onlyQwen 3 4B
8 GB VRAMQwen 3 8B
12 to 16 GB VRAMQwen 3 14B
24 GB VRAMQwen 3 32B or Qwen 3 30B-A3B
Large Mac or multi-GPUQwen 3 235B-A22B

Check the VRAM column above against your card and leave some room for a longer context.

Dense or MoE at 30B? Qwen 3 32B is a dense model. Qwen 3 30B-A3B is a mixture-of-experts model that computes only a few billion parameters per token, so it is much faster at similar memory. Try both; see MoE Models and VRAM.

How memory changes with quantization and context

For Qwen 3 14B:

QuantizationVRAM at 4K contextVRAM at 8K contextVRAM at 32K context
Q3_K_M8.4 GB9.0 GB12.8 GB
Q4_K_M10.0 GB10.6 GB14.4 GB
Q5_K_M11.4 GB12.1 GB15.8 GB
Q6_K12.9 GB13.5 GB17.3 GB
Q8_016.2 GB16.9 GB20.6 GB

Run it with Ollama

  1. Install Ollama from ollama.com for macOS, Windows or Linux.
  2. Run the size you chose. For the 8B model:

ollama run qwen3:8b

For the 14B model:

ollama run qwen3:14b

The first run downloads the weights.

  1. To use a longer context in a session:

/set parameter num_ctx 8192

  1. Ollama also serves a local API on port 11434 that other apps can use.

Thinking mode

Qwen 3 models can produce a reasoning section before the answer. It helps on math, logic and difficult code, and costs time. For quick questions, turn it off; Qwen 3 accepts a /no_think instruction in the prompt, and recent versions of local tools expose a switch for it.

Settings worth knowing

  • Quantization. Q4_K_M is the default choice. Use Q5 or Q6 if you have spare memory and want a little more quality.
  • Context length. Every extra token of context uses memory. Start at 8K.
  • Temperature. Lower for code and factual answers, higher for creative writing.

What Qwen 3 is good at

Qwen models are strong for their size in coding and in multilingual work, including Indian languages; see Best Open-Source LLM for Hindi. For how specific sizes rank on each task, open the LLM Leaderboard and switch on the self-host filter.

Newer Qwen releases

Alibaba releases new Qwen versions frequently. The leaderboard lists the versions currently tracked and which have open weights. The steps on this page are the same for any of them: check the memory, pick the tag, run.

Frequently asked questions

Which Qwen 3 model fits in 8 GB of VRAM?

Qwen 3 8B needs about 6.5 GB at Q4_K_M with an 8K context, and Qwen 3 4B needs about 4.2 GB.

Which Qwen 3 model is best for a 24 GB GPU?

Qwen 3 32B needs about 21.2 GB at Q4_K_M with an 8K context, which fits a 24 GB card.

What is the Ollama command for Qwen 3?

ollama run followed by the tag for the size you want, for example ollama run qwen3:8b.

What is the thinking mode in Qwen 3?

Qwen 3 models can reason step by step before answering. It improves hard questions and makes replies slower. It can be switched off for simple requests.

What licence is Qwen 3 under?

The Qwen 3 open-weight models were released under the Apache 2.0 licence. Confirm on each model's Hugging Face page.

Related