Guide

Best Local LLM for Agentic Coding in 2026: What Works on Your Own Hardware

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

A coding agent reads files, runs commands and edits code over many steps. Doing that with a model on your own machine keeps your code private and costs nothing per token, but it asks more of the model than chat does. This guide covers which open-weight models lead and what hardware they need.

Switch on the self-host filter to see open-weight models only.

See the agents ranking →

What agentic coding asks of a model

In chat, a model answers once. An agent works in a loop: read a file, decide, call a tool, read the result, decide again. That needs four things:

  1. Reliable tool calls. Every call must be in exactly the right format. One malformed call stops the loop.
  2. Clean edits. Changes have to apply to the file as written. See What is Aider Polyglot?.
  3. Memory of the task. After twenty steps the model still has to remember what it was doing.
  4. Recovery. When a command fails, it should read the error and try something different.

Small models manage the first step and fall down on the third and fourth.

Leading open-weight models for agents

#ModelProviderGradePrice in / out per 1MContext
1Kimi K3Moonshot AIS$3 / $151M
2DeepSeek V4.1 FlashDeepSeekS$0.15 / $0.61M
3Kimi K2.6Moonshot AIS$0.95 / $4256K
4DeepSeek V4-ProDeepSeekS$0.435 / $0.871M
5MiMo-V2.5-ProXiaomiS$0.435 / $0.871M
6Qwen3.6 35B-A3BAlibabaS$0.38 / $2.25262K
7Qwen3.6 27BAlibabaS—262K
8Qwen 3 235B-A22BAlibabaA—41K

This is the id8 agents ranking limited to open-weight models. It uses SWE-bench Verified, Arena Elo for web development and Aider Polyglot.

For coding in general

#ModelProviderGradePrice in / out per 1MContext
1DeepSeek V4-FlashDeepSeekS$0.14 / $0.281M
2Kimi K2.6Moonshot AIS$0.95 / $4256K
3DeepSeek V4-ProDeepSeekS$0.435 / $0.871M
4Kimi K3Moonshot AIS$3 / $151M
5Kimi K2.5Moonshot AIS$0.6 / $3262K
6GLM-5.2Zhipu AIS$1.4 / $4.41M
7GLM-5.3Zhipu AIS$1.4 / $4.41M
8DeepSeek V4.1 FlashDeepSeekS$0.15 / $0.61M

What will actually run on your machine

The strongest open-weight coding models are large. Before choosing, check the memory:

  • Open the model in Can I Run LLM? and set the context to at least 32K tokens, since that is what an agent needs.
  • For what fits in 8 to 32 GB of VRAM, see Best Local LLM by VRAM. Remember that those tables use an 8K context; a 32K context needs more.

As an example of how context changes the requirement, here is Qwen 3 32B:

QuantizationVRAM at 4K contextVRAM at 8K contextVRAM at 32K context
Q3_K_M16.7 GB17.7 GB23.7 GB
Q4_K_M20.2 GB21.2 GB27.2 GB
Q5_K_M23.2 GB24.2 GB30.2 GB
Q6_K26.5 GB27.5 GB33.5 GB
Q8_033.7 GB34.7 GB40.7 GB

On a 24 GB card, a model in the 30B class at Q4 leaves limited room for context. That trade, a better model or a longer context, is the central decision for local agents.

Realistic expectations by hardware

HardwareWhat works
8 to 12 GB VRAMCode completion and single-file edits. Multi-step agents are unreliable.
16 to 24 GB VRAMShort agent tasks: a function, a test, a small refactor in a few files.
32 to 48 GB, or a large MacLonger tasks with a 30B-class model and a real context window.
Multi-GPU or serverThe largest open-weight models, close to API quality on coding benchmarks.

Setting up a local coding agent

  1. Serve the model with Ollama or LM Studio. Both expose an OpenAI-compatible local endpoint. See Ollama vs LM Studio.
  2. Raise the context length. The default in local tools is too small for agents. Set it explicitly.
  3. Point your coding tool at the local endpoint. Most open-source coding agents and editor extensions accept a custom base URL.
  4. Choose a model with tool-calling support. Check the model card. Without it the agent has to parse free text, which is fragile.
  5. Start with small, well-defined tasks and commit your work first, so any change can be undone.

Habits that make local agents more reliable

  • One task per session. A fresh context is the cheapest fix for a confused agent.
  • Give it the file names. Do not make a small model search the repository.
  • Have tests it can run. A failing test is a clear signal; "it looks right" is not.
  • Review every diff. Local or not, the agent's output is a draft.
  • Stop loops early. If it repeats a failed command twice, intervene.

When to use an API instead

If the work is not confidential and the task is long or spans many files, a strong API model will finish faster and with fewer errors. Many developers use a local model for quick private edits and an API for large jobs. The cheapest model graded A or better for agents is Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens); the full list is in Best LLM for Agents.

Frequently asked questions

What is the best open-weight LLM for agentic coding?

On the id8 agents ranking as of October 2, 2026, the leading open-weight models are Kimi K3, DeepSeek V4.1 Flash and Kimi K2.6.

Can a local model match a flagship API model for agents?

The best open-weight models are close on some coding benchmarks, but they are large. Models that fit on a single consumer GPU are noticeably weaker at long multi-step tasks than flagship APIs.

How much context does a coding agent need?

More than chat. File contents and command output accumulate, so 32K tokens is a practical minimum and more is better. Context uses memory, so check it for your model.

Why do local agents get stuck in loops?

Smaller models lose track of earlier steps, repeat failed commands, or produce edits that do not apply. Shorter tasks, a larger context and a model with good tool-calling support reduce this.

What is the cheapest API if local is not good enough?

The cheapest model graded A or better for agents is Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens).

Related