Local AI Hardware

How to Run LLMs on Apple Silicon Mac (M1, M2, M3) — Complete Guide

Last updated: 2026-03-12

Apple Silicon is one of the best local AI platforms in 2026 because CPU and GPU share one unified memory pool. That changes what "fits" compared to discrete-GPU PCs. A Mac with enough unified memory can run models that would fail on low-VRAM cards, while remaining quiet and stable for long sessions. This guide explains sizing, setup, and practical performance choices for M1, M2, and M3.

🛠️ Developer's Note

Apple Silicon is my primary local AI platform, and I think it's underrated for this use case. Unified memory means the GPU can access the full memory pool — no PCIe bottleneck, no GPU VRAM ceiling. This guide reflects my actual daily experience running models on M-series Macs.

Why Apple Silicon Is Different for LLMs

On most desktop PCs, VRAM is separate from system RAM. If a model needs 14GB and your GPU has 8GB, you are forced into heavy offloading. Apple Silicon uses unified memory, so CPU and GPU draw from the same pool. This means a 16GB or 32GB Mac can run model sizes that look impossible if you only think in discrete VRAM terms.

This does not mean unlimited performance. Memory bandwidth, thermals, and model architecture still matter. But unified memory gives Apple laptops and desktops a predictable local-inference advantage, especially for medium-size quantized models.

Unified Memory Explained

Unified memory is a shared pool used by CPU, GPU, and neural processing units. There is no separate copy between system memory and VRAM, which reduces overhead and allows larger contiguous allocations for local inference. In practical terms, if your Mac has enough free unified memory after macOS and active apps, more of that pool can be used for model weights and KV cache.

That is why an M1 Pro 16GB can run models in the 7B-8B class comfortably at Q4, and why higher-tier machines scale much further.

What Can Each Mac Run?

Chip Tier Unified Memory Max Model Size (Q4_K_M) Example Models
M1/M2 base8GB~5.5GB usablePhi-3 Mini, Llama 3.2 3B
M1/M216GB~12GBLlama 3.1 8B, Mistral 7B
M1/M2 Pro32GB~26GB32B models, partial 70B Q4
M2/M3 Ultra192GBVery large models70B Q8 class workloads

Step-by-step: Install Ollama on Mac

If you use Homebrew, install is one command. Then pull a starter model and run it directly in the terminal.

brew install ollama
ollama pull llama3.2
ollama run llama3.2

Ollama uses Metal acceleration automatically on Apple Silicon, so no CUDA setup is needed.

Performance Tips for Apple Silicon

  • Use Q4_K_M as default for better memory efficiency.
  • Use Q8_0 only when memory headroom exists (typically 32GB+ machines).
  • Close heavy apps before long runs to free unified memory.
  • Keep context size aligned with task needs to reduce KV cache usage.
  • Prefer smaller specialized models for latency-sensitive workflows.

These changes often improve perceived speed more than raw benchmark chasing.

Check Your Mac's Model Limits

Use Can I Run LLM and select Apple Silicon profile to validate fit before downloading large models. It is faster than trial-and-error installs and avoids repeated pull/remove cycles.

Check your GPU before you download

Can I Run LLM — Free VRAM Calculator →

Apple Silicon can be an excellent local AI stack when expectations match memory tier. Start with a model that fits cleanly, confirm performance, then scale. That workflow gives better results than forcing oversized models and debugging avoidable memory failures.

Model Planning for Mac Users

The easiest way to stay productive is to pin two models: one fast daily model and one deeper backup model. For example, on 16GB machines you might keep Mistral 7B as your daily model and an 8B reasoning model for harder prompts. On 32GB systems, you can keep a 14B-class model for analysis tasks while still retaining a smaller fast model for quick iterative work.

This two-model strategy avoids constant pull/remove churn and helps maintain consistent behavior across your tools. It also gives better battery control when you switch between light and heavy workloads.

Thermal and Battery Considerations

Apple Silicon laptops can sustain long inference sessions, but sustained high-load generation still affects battery and thermals. If you run local models unplugged, reduce context length and use lower quantization to keep energy usage manageable. For long coding or research sessions, run on power and close unused apps to maximize available memory bandwidth.

These operational choices matter more than benchmark numbers. Real-world local AI experience on Mac is usually defined by consistency, not peak throughput.

For teams standardizing on Mac hardware, document one recommended model set per memory tier (16GB, 32GB, 64GB+). This avoids repeated onboarding confusion and makes local assistant behavior more predictable across contributors.

Keep one fallback lightweight model installed as well. During travel or battery-constrained work, switching to a smaller model can preserve responsiveness without interrupting your local workflow.

That single fallback habit prevents most mid-session performance surprises.

Combined with a clear memory-tier playbook, Apple Silicon becomes one of the most predictable local AI environments available to individual developers and small teams.

Limitations / When NOT to Use This

  • Unified memory is shared with the OS and all running apps — actual available memory for LLM inference is typically 60-75% of total installed RAM, not the full amount
  • Token generation speed on Apple Silicon is competitive but generally 20-40% slower than equivalent NVIDIA GPUs for the same model size and quantization
  • macOS memory pressure can cause the system to swap when running large models, leading to sudden slowdowns — monitor Activity Monitor's memory tab during inference
  • Not all inference runtimes are equally optimized for Metal — Ollama and llama.cpp have good Metal support, but some less common runtimes may fall back to CPU-only execution

Related Guides