Hardware Guide
RTX 5090 vs RTX 4090 vs M4 Max for Local LLMs: What Each Can Run
Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset
For local LLMs, memory decides what you can run and memory bandwidth decides how fast. The RTX 5090, the RTX 4090 and an M4 Max MacBook take three different positions on that trade. This page shows which models fit on each.
Pick a model and context length; see every GPU and Mac that runs it.
Find the cheapest GPU for your model →The three machines
| RTX 4090 | RTX 5090 | M4 Max MacBook Pro | |
|---|---|---|---|
| Memory for the model | 24 GB VRAM | 32 GB VRAM | About 75% of unified memory |
| Memory options | Fixed | Fixed | 36 GB to 128 GB total |
| Needs | A desktop PC with a strong power supply | The same, with more power | Nothing else |
| Software | CUDA: everything works | CUDA: everything works | Metal and MLX: most local tools work |
| Fine-tuning | Well supported | Well supported | Supported through MLX |
What fits on an RTX 4090 (24 GB)
Largest models that fit at Q4_K_M with an 8K context:
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Qwen3.6 35B-A3B | Alibaba | 35B | 21.2 GB | qwen3.6:35b |
| QwQ 32B | Alibaba | 32.5B | 21.5 GB | qwq:32b |
| Qwen 2.5 32B | Alibaba | 32B | 21.2 GB | qwen2.5:32b |
| Qwen 2.5-Coder 32B | Alibaba | 32B | 21.2 GB | qwen2.5-coder:32b |
| Qwen 3 32B | Alibaba | 32B | 21.2 GB | qwen3:32b |
| DeepSeek R1 Distill 32B | DeepSeek | 32B | 21.2 GB | deepseek-r1:32b |
What fits on an RTX 5090 (32 GB)
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Mixtral 8x7B | Mistral AI | 46.7B | 28.7 GB | mixtral:8x7b |
| Qwen3.6 35B-A3B | Alibaba | 35B | 21.2 GB | qwen3.6:35b |
| QwQ 32B | Alibaba | 32.5B | 21.5 GB | qwq:32b |
| Qwen 2.5 32B | Alibaba | 32B | 21.2 GB | qwen2.5:32b |
| Qwen 2.5-Coder 32B | Alibaba | 32B | 21.2 GB | qwen2.5-coder:32b |
| Qwen 3 32B | Alibaba | 32B | 21.2 GB | qwen3:32b |
Compare the two tables. If the same models top both, the extra 8 GB is buying you a longer context or a higher-quality quantization, not a bigger model.
What fits on a 64 GB M4 Max (about 48 GB usable)
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Qwen 2.5 72B | Alibaba | 72B | 44.6 GB | qwen2.5:72b |
| Llama 3.3 70B | Meta | 70B | 43.5 GB | llama3.3:70b |
| DeepSeek R1 Distill 70B | DeepSeek | 70B | 43.5 GB | deepseek-r1:70b |
| Mixtral 8x7B | Mistral AI | 46.7B | 28.7 GB | mixtral:8x7b |
| Qwen3.6 35B-A3B | Alibaba | 35B | 21.2 GB | qwen3.6:35b |
| QwQ 32B | Alibaba | 32.5B | 21.5 GB | qwq:32b |
What fits on a 128 GB M4 Max (about 96 GB usable)
| Model | Provider | Parameters | VRAM at Q4_K_M, 8K context | Ollama tag |
|---|---|---|---|---|
| Qwen3.5 122B-A10B | Alibaba | 122B | 71.0 GB | qwen3.5:122b |
| GPT-oss 120B | OpenAI | 117B | 68.3 GB | — |
| Llama 4 Scout | Meta | 109B | 64.9 GB | llama4:scout |
| Llama 3.2 Vision 90B | Meta | 90B | 55.8 GB | llama3.2-vision:90b |
| Qwen3 Coder Next | Alibaba | 80B | 47.1 GB | — |
| Qwen 2.5 72B | Alibaba | 72B | 44.6 GB | qwen2.5:72b |
This is where a Mac does something no single consumer GPU can: hold a very large model entirely in fast memory.
Speed
Token generation speed depends mainly on how fast the hardware can read the model's weights, that is, on memory bandwidth, and on compute.
- The RTX 5090 is the fastest of the three for any model that fits in its 32 GB.
- The RTX 4090 is somewhat slower and still very fast.
- The M4 Max is slower per token than both NVIDIA cards on models that fit on them, and usable for interactive work on mid-size models.
For a model that does not fit in VRAM on a PC, part of it runs from system RAM and speed collapses. On such a model a Mac with enough memory is far faster than a GPU that has to offload.
Exact tokens-per-second figures vary with the model, quantization, context and software version, so this page does not quote them. Can I Run LLM? gives an estimate for your combination.
How to decide
Choose the RTX 4090 if the models you want are in the 24 GB table. It is the best-supported card for local AI, and older stock or used cards cost less than a 5090.
Choose the RTX 5090 if you want the fastest single-GPU setup, need a longer context with 30B-class models, or also plan to fine-tune.
Choose the M4 Max if you want one portable machine, want to run models larger than 32 GB, or value silence and low power draw. Buy as much memory as you can; it cannot be upgraded later.
Consider two used 24 GB cards if you want 48 GB on a PC and are comfortable with a larger build.
Things people overlook
- Power and heat. A 5090 system needs a high-wattage power supply and good airflow. A Mac runs the same model on a fraction of the power.
- Context length. A long context can add several gigabytes. Budget for it; see What is the KV cache?.
- Software. Some training libraries and newer features are CUDA-first. Check that your tools support Apple Silicon before choosing a Mac for fine-tuning.
- Price in India. Import duty and availability move GPU prices more than list prices suggest. Compare current local prices before deciding.
Check a specific model
Pick the model you care about in What GPU Do I Need? and it lists every GPU and Mac configuration that runs it at your context length.
Frequently asked questions
Is the RTX 5090 worth it over the RTX 4090 for LLMs?
The 5090 has 32 GB of VRAM against the 4090's 24 GB. If the model you want fits in 24 GB, the 4090 is enough. The extra 8 GB matters when it lets you run the next model size or a much longer context.
Can an M4 Max run larger models than an RTX 5090?
Yes, when configured with more memory. A Mac can use about three quarters of its unified memory for the model, so a 64 GB M4 Max has roughly 48 GB available and a 128 GB one roughly 96 GB.
Which is fastest?
For models that fit in VRAM, the NVIDIA cards generate tokens faster than Apple Silicon. A Mac's advantage is running models too large for any single consumer GPU.
What about two GPUs?
Two 24 GB cards give 48 GB for inference with tools that split a model across GPUs. It needs a suitable motherboard, power supply and cooling.