Hardware Guide

RTX 5090 vs RTX 4090 vs M4 Max for Local LLMs: What Each Can Run

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

For local LLMs, memory decides what you can run and memory bandwidth decides how fast. The RTX 5090, the RTX 4090 and an M4 Max MacBook take three different positions on that trade. This page shows which models fit on each.

Pick a model and context length; see every GPU and Mac that runs it.

Find the cheapest GPU for your model →

The three machines

RTX 4090RTX 5090M4 Max MacBook Pro
Memory for the model24 GB VRAM32 GB VRAMAbout 75% of unified memory
Memory optionsFixedFixed36 GB to 128 GB total
NeedsA desktop PC with a strong power supplyThe same, with more powerNothing else
SoftwareCUDA: everything worksCUDA: everything worksMetal and MLX: most local tools work
Fine-tuningWell supportedWell supportedSupported through MLX

What fits on an RTX 4090 (24 GB)

Largest models that fit at Q4_K_M with an 8K context:

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Qwen3.6 35B-A3BAlibaba35B21.2 GBqwen3.6:35b
QwQ 32BAlibaba32.5B21.5 GBqwq:32b
Qwen 2.5 32BAlibaba32B21.2 GBqwen2.5:32b
Qwen 2.5-Coder 32BAlibaba32B21.2 GBqwen2.5-coder:32b
Qwen 3 32BAlibaba32B21.2 GBqwen3:32b
DeepSeek R1 Distill 32BDeepSeek32B21.2 GBdeepseek-r1:32b

What fits on an RTX 5090 (32 GB)

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Mixtral 8x7BMistral AI46.7B28.7 GBmixtral:8x7b
Qwen3.6 35B-A3BAlibaba35B21.2 GBqwen3.6:35b
QwQ 32BAlibaba32.5B21.5 GBqwq:32b
Qwen 2.5 32BAlibaba32B21.2 GBqwen2.5:32b
Qwen 2.5-Coder 32BAlibaba32B21.2 GBqwen2.5-coder:32b
Qwen 3 32BAlibaba32B21.2 GBqwen3:32b

Compare the two tables. If the same models top both, the extra 8 GB is buying you a longer context or a higher-quality quantization, not a bigger model.

What fits on a 64 GB M4 Max (about 48 GB usable)

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Qwen 2.5 72BAlibaba72B44.6 GBqwen2.5:72b
Llama 3.3 70BMeta70B43.5 GBllama3.3:70b
DeepSeek R1 Distill 70BDeepSeek70B43.5 GBdeepseek-r1:70b
Mixtral 8x7BMistral AI46.7B28.7 GBmixtral:8x7b
Qwen3.6 35B-A3BAlibaba35B21.2 GBqwen3.6:35b
QwQ 32BAlibaba32.5B21.5 GBqwq:32b

What fits on a 128 GB M4 Max (about 96 GB usable)

ModelProviderParametersVRAM at Q4_K_M, 8K contextOllama tag
Qwen3.5 122B-A10BAlibaba122B71.0 GBqwen3.5:122b
GPT-oss 120BOpenAI117B68.3 GB—
Llama 4 ScoutMeta109B64.9 GBllama4:scout
Llama 3.2 Vision 90BMeta90B55.8 GBllama3.2-vision:90b
Qwen3 Coder NextAlibaba80B47.1 GB—
Qwen 2.5 72BAlibaba72B44.6 GBqwen2.5:72b

This is where a Mac does something no single consumer GPU can: hold a very large model entirely in fast memory.

Speed

Token generation speed depends mainly on how fast the hardware can read the model's weights, that is, on memory bandwidth, and on compute.

  • The RTX 5090 is the fastest of the three for any model that fits in its 32 GB.
  • The RTX 4090 is somewhat slower and still very fast.
  • The M4 Max is slower per token than both NVIDIA cards on models that fit on them, and usable for interactive work on mid-size models.

For a model that does not fit in VRAM on a PC, part of it runs from system RAM and speed collapses. On such a model a Mac with enough memory is far faster than a GPU that has to offload.

Exact tokens-per-second figures vary with the model, quantization, context and software version, so this page does not quote them. Can I Run LLM? gives an estimate for your combination.

How to decide

Choose the RTX 4090 if the models you want are in the 24 GB table. It is the best-supported card for local AI, and older stock or used cards cost less than a 5090.

Choose the RTX 5090 if you want the fastest single-GPU setup, need a longer context with 30B-class models, or also plan to fine-tune.

Choose the M4 Max if you want one portable machine, want to run models larger than 32 GB, or value silence and low power draw. Buy as much memory as you can; it cannot be upgraded later.

Consider two used 24 GB cards if you want 48 GB on a PC and are comfortable with a larger build.

Things people overlook

  • Power and heat. A 5090 system needs a high-wattage power supply and good airflow. A Mac runs the same model on a fraction of the power.
  • Context length. A long context can add several gigabytes. Budget for it; see What is the KV cache?.
  • Software. Some training libraries and newer features are CUDA-first. Check that your tools support Apple Silicon before choosing a Mac for fine-tuning.
  • Price in India. Import duty and availability move GPU prices more than list prices suggest. Compare current local prices before deciding.

Check a specific model

Pick the model you care about in What GPU Do I Need? and it lists every GPU and Mac configuration that runs it at your context length.

Frequently asked questions

Is the RTX 5090 worth it over the RTX 4090 for LLMs?

The 5090 has 32 GB of VRAM against the 4090's 24 GB. If the model you want fits in 24 GB, the 4090 is enough. The extra 8 GB matters when it lets you run the next model size or a much longer context.

Can an M4 Max run larger models than an RTX 5090?

Yes, when configured with more memory. A Mac can use about three quarters of its unified memory for the model, so a 64 GB M4 Max has roughly 48 GB available and a 128 GB one roughly 96 GB.

Which is fastest?

For models that fit in VRAM, the NVIDIA cards generate tokens faster than Apple Silicon. A Mac's advantage is running models too large for any single consumer GPU.

What about two GPUs?

Two 24 GB cards give 48 GB for inference with tools that split a model across GPUs. It needs a suitable motherboard, power supply and cooling.

Related