Guide
Best Local LLM for Agentic Coding in 2026: What Works on Your Own Hardware
Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset
A coding agent reads files, runs commands and edits code over many steps. Doing that with a model on your own machine keeps your code private and costs nothing per token, but it asks more of the model than chat does. This guide covers which open-weight models lead and what hardware they need.
Switch on the self-host filter to see open-weight models only.
See the agents ranking →What agentic coding asks of a model
In chat, a model answers once. An agent works in a loop: read a file, decide, call a tool, read the result, decide again. That needs four things:
- Reliable tool calls. Every call must be in exactly the right format. One malformed call stops the loop.
- Clean edits. Changes have to apply to the file as written. See What is Aider Polyglot?.
- Memory of the task. After twenty steps the model still has to remember what it was doing.
- Recovery. When a command fails, it should read the error and try something different.
Small models manage the first step and fall down on the third and fourth.
Leading open-weight models for agents
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | Kimi K3 | Moonshot AI | S | $3 / $15 | 1M |
| 2 | DeepSeek V4.1 Flash | DeepSeek | S | $0.15 / $0.6 | 1M |
| 3 | Kimi K2.6 | Moonshot AI | S | $0.95 / $4 | 256K |
| 4 | DeepSeek V4-Pro | DeepSeek | S | $0.435 / $0.87 | 1M |
| 5 | MiMo-V2.5-Pro | Xiaomi | S | $0.435 / $0.87 | 1M |
| 6 | Qwen3.6 35B-A3B | Alibaba | S | $0.38 / $2.25 | 262K |
| 7 | Qwen3.6 27B | Alibaba | S | — | 262K |
| 8 | Qwen 3 235B-A22B | Alibaba | A | — | 41K |
This is the id8 agents ranking limited to open-weight models. It uses SWE-bench Verified, Arena Elo for web development and Aider Polyglot.
For coding in general
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | DeepSeek V4-Flash | DeepSeek | S | $0.14 / $0.28 | 1M |
| 2 | Kimi K2.6 | Moonshot AI | S | $0.95 / $4 | 256K |
| 3 | DeepSeek V4-Pro | DeepSeek | S | $0.435 / $0.87 | 1M |
| 4 | Kimi K3 | Moonshot AI | S | $3 / $15 | 1M |
| 5 | Kimi K2.5 | Moonshot AI | S | $0.6 / $3 | 262K |
| 6 | GLM-5.2 | Zhipu AI | S | $1.4 / $4.4 | 1M |
| 7 | GLM-5.3 | Zhipu AI | S | $1.4 / $4.4 | 1M |
| 8 | DeepSeek V4.1 Flash | DeepSeek | S | $0.15 / $0.6 | 1M |
What will actually run on your machine
The strongest open-weight coding models are large. Before choosing, check the memory:
- Open the model in Can I Run LLM? and set the context to at least 32K tokens, since that is what an agent needs.
- For what fits in 8 to 32 GB of VRAM, see Best Local LLM by VRAM. Remember that those tables use an 8K context; a 32K context needs more.
As an example of how context changes the requirement, here is Qwen 3 32B:
| Quantization | VRAM at 4K context | VRAM at 8K context | VRAM at 32K context |
|---|---|---|---|
| Q3_K_M | 16.7 GB | 17.7 GB | 23.7 GB |
| Q4_K_M | 20.2 GB | 21.2 GB | 27.2 GB |
| Q5_K_M | 23.2 GB | 24.2 GB | 30.2 GB |
| Q6_K | 26.5 GB | 27.5 GB | 33.5 GB |
| Q8_0 | 33.7 GB | 34.7 GB | 40.7 GB |
On a 24 GB card, a model in the 30B class at Q4 leaves limited room for context. That trade, a better model or a longer context, is the central decision for local agents.
Realistic expectations by hardware
| Hardware | What works |
|---|---|
| 8 to 12 GB VRAM | Code completion and single-file edits. Multi-step agents are unreliable. |
| 16 to 24 GB VRAM | Short agent tasks: a function, a test, a small refactor in a few files. |
| 32 to 48 GB, or a large Mac | Longer tasks with a 30B-class model and a real context window. |
| Multi-GPU or server | The largest open-weight models, close to API quality on coding benchmarks. |
Setting up a local coding agent
- Serve the model with Ollama or LM Studio. Both expose an OpenAI-compatible local endpoint. See Ollama vs LM Studio.
- Raise the context length. The default in local tools is too small for agents. Set it explicitly.
- Point your coding tool at the local endpoint. Most open-source coding agents and editor extensions accept a custom base URL.
- Choose a model with tool-calling support. Check the model card. Without it the agent has to parse free text, which is fragile.
- Start with small, well-defined tasks and commit your work first, so any change can be undone.
Habits that make local agents more reliable
- One task per session. A fresh context is the cheapest fix for a confused agent.
- Give it the file names. Do not make a small model search the repository.
- Have tests it can run. A failing test is a clear signal; "it looks right" is not.
- Review every diff. Local or not, the agent's output is a draft.
- Stop loops early. If it repeats a failed command twice, intervene.
When to use an API instead
If the work is not confidential and the task is long or spans many files, a strong API model will finish faster and with fewer errors. Many developers use a local model for quick private edits and an API for large jobs. The cheapest model graded A or better for agents is Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens); the full list is in Best LLM for Agents.
Frequently asked questions
What is the best open-weight LLM for agentic coding?
On the id8 agents ranking as of October 2, 2026, the leading open-weight models are Kimi K3, DeepSeek V4.1 Flash and Kimi K2.6.
Can a local model match a flagship API model for agents?
The best open-weight models are close on some coding benchmarks, but they are large. Models that fit on a single consumer GPU are noticeably weaker at long multi-step tasks than flagship APIs.
How much context does a coding agent need?
More than chat. File contents and command output accumulate, so 32K tokens is a practical minimum and more is better. Context uses memory, so check it for your model.
Why do local agents get stuck in loops?
Smaller models lose track of earlier steps, repeat failed commands, or produce edits that do not apply. Shorter tasks, a larger context and a model with good tool-calling support reduce this.
What is the cheapest API if local is not good enough?
The cheapest model graded A or better for agents is Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens).