Local AI Hardware

How to Run Any LLM Locally with Ollama — Complete Step-by-Step Guide (2026)

Last updated: 2026-05-07 — added Gemma 4, Phi-4-reasoning-plus, Mistral Small 3.2

Ollama is the simplest way to run open models locally. It gives you one-command installs, automatic hardware detection, model management, and a local API compatible with common app patterns. Instead of manually wiring runtimes and model files, you install once and start pulling models immediately. This guide covers setup on Mac, Windows, Linux, plus practical commands and failure fixes.

🛠️ Developer's Note

Ollama made local LLM deployment dramatically simpler than llama.cpp used to be. I recommend it to anyone starting out because the setup time dropped from 'hours of compilation' to 'one terminal command'. This guide covers the practical steps I walk people through most often.

Prerequisites

  • GPU with about 6GB+ VRAM, or Apple Silicon Mac with sufficient unified memory.
  • NVIDIA drivers installed on Windows/Linux (for GPU acceleration).
  • macOS 13.1+ on Apple Silicon devices.
  • At least 20GB free disk space for model downloads and cache.

You can run Ollama on CPU-only systems, but inference speed usually drops significantly for larger models.

Install Ollama

Use the command path that matches your operating system.

# macOS (Homebrew)
brew install ollama

# Linux
curl -fsSL https://ollama.com/install.sh | sh

# Windows
# Download and run the .exe installer from ollama.com

After installation, the local service should be available. On some systems you may need to run ollama serve manually.

Pull and Run Your First Model

Start with a small model to validate your environment.

ollama pull llama3.2
ollama run llama3.2

ollama pull downloads and stores model artifacts locally (typically under ~/.ollama/models). ollama run starts an interactive session.

Popular Models and Commands

Model Command Min VRAM
Gemma 4 E2Bollama pull gemma4:e2b2.5GB
Gemma 4 E4Bollama pull gemma4:e4b3GB
Llama 3.2 3Bollama pull llama3.2:3b3GB
Llama 3.1 8Bollama pull llama3.1:8b5GB
Mistral 7Bollama pull mistral4.5GB
Mistral Small 3.2 24Bollama pull mistral-small3.2:24b14GB
DeepSeek R1 7Bollama pull deepseek-r1:7b5GB
Phi-3.5 Miniollama pull phi3.53.8GB
Phi-4-reasoning-plus 14Bollama pull phi4-reasoning:plus9GB
Qwen 2.5 14Bollama pull qwen2.5:14b9GB

Use the Ollama API

Ollama exposes a local endpoint at http://localhost:11434/v1. You can use it similarly to OpenAI-style clients.

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "messages": [{"role": "user", "content": "Summarize this repo in 5 bullet points."}]
  }'
import requests

resp = requests.post(
    "http://localhost:11434/v1/chat/completions",
    json={
        "model": "llama3.2",
        "messages": [{"role": "user", "content": "Explain quantization in simple terms."}]
    },
    timeout=60,
)
print(resp.json())

Switch Between Models

ollama list
ollama run mistral
ollama rm llama3.2:3b

Use ollama rm to reclaim disk if you are testing many models quickly.

Check if Your GPU Is Supported

Model pulls can take time, so validate hardware first with Can I Run LLM.

Check your GPU before you download

Can I Run LLM — Free VRAM Calculator →

Common Errors and Fixes

Error Cause Fix
could not connect to ollamaService not runningRun ollama serve
Out of memory / crashModel too largeUse smaller model or lower quantization
Very slow inferenceGPU not detectedCheck drivers/CUDA setup

A Practical First-Week Ollama Plan

If you are new to local inference, do not install ten models on day one. Keep a structured first week: one small fast model for interaction, one medium model for quality checks, and one benchmark script to compare response stability. This prevents configuration drift and makes performance tuning straightforward.

  1. Day 1: install and run one small model end-to-end.
  2. Day 2: add one medium model and compare speed vs quality.
  3. Day 3: set up a minimal API client in your app or script.
  4. Day 4: test prompt templates and output format controls.
  5. Day 5: clean unused models and lock your stable pair.

By the end of that cycle, you have a reproducible local setup rather than an unmaintainable model collection.

Security and Operational Notes

Ollama runs a local service endpoint by default. On personal machines this is usually fine, but on shared or networked environments you should review firewall rules and service exposure. Treat local model endpoints like any internal service: restrict unnecessary access and avoid binding publicly unless you intentionally proxy and secure it.

Also maintain disk hygiene. Large model pulls can consume tens of gigabytes quickly. Use ollama list weekly, remove inactive models, and keep one rollback model version if a newer tag changes output behavior.

Performance Tuning Checklist

  • Use Q4 or Q8 based on available memory, not preference alone.
  • Keep context size only as large as your task requires.
  • Close GPU-heavy applications before long inference sessions.
  • Benchmark with fixed prompts after any model or driver update.
  • Track tokens/second over time to catch regressions early.

These five checks usually solve most \"Ollama feels slow\" reports without hardware upgrades.

Using Ollama in Daily Dev Work

The biggest productivity gain comes when Ollama is integrated into your normal tooling. Keep one terminal session pinned for interactive prompting, one API script for batch tasks, and one editor shortcut for common transformations such as test generation or commit-message drafts. This makes local AI a stable part of workflow rather than an occasional experiment.

If outputs drift, roll back to a known model tag and re-run your benchmark prompts. Version control for model tags and prompt templates is as important as code versioning when teams depend on repeatable assistant behavior.

Store your best-performing model tags in project docs so collaborators can reproduce the same local assistant behavior without trial-and-error setup.

In team environments, pair model tags with prompt-template versions to keep output style stable.

That extra discipline turns Ollama from a personal experiment into a dependable local inference platform that can support real engineering workflows.

Small documentation habits create big operational stability over time.

Ollama works best when you treat model selection as a hardware fit problem, not just a leaderboard problem. Start from memory limits, then choose the strongest model that runs reliably.

Limitations / When NOT to Use This

  • Ollama abstracts away quantization choices — you get the default quant the model publisher provides, which may not be optimal for your hardware
  • Model management in Ollama can use significant disk space — each model variant is stored separately, and a few models can quickly fill 50-100 GB of storage
  • Ollama's API server runs on localhost by default, which is fine for personal use but requires additional configuration (and security review) for network-accessible deployments
  • Performance tuning options are limited compared to raw llama.cpp — advanced users who need custom batch sizes, thread pinning, or GPU layer splitting may outgrow Ollama

Related Guides