Monthly Update
LLM Leaderboard Update, October 2026: Who Leads Each Task Now
Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset
A plain summary of where the id8 LLM Leaderboard stands as of October 1, 2026: the leaders overall and by task, the strongest open-weight models, and the cheapest models that still grade A or better. The tables on this page are generated from the same data as the leaderboard.
All models, every benchmark, prices and context windows. Free, no signup.
Open the full leaderboard →The short version
- Overall: Claude Fable 5, Gemini 3.1 Pro and Claude Opus 4.8 lead.
- Coding: Claude Opus 5.5, GPT-6 Astra and Claude Fable 5.
- Reasoning: Gemini 3.1 Pro, GPT-6 Astra and Claude Fable 5.
- Best open-weight overall: DeepSeek V4-Pro, Kimi K2.5 and Kimi K3.
- Cheapest model graded A or better overall: Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens).
The leaderboard currently tracks 164 models, of which 95 can be run on your own hardware.
Top 10 overall
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | S | $10 / $50 | 1M |
| 2 | Gemini 3.1 Pro | S | $2 / $12 | 1M | |
| 3 | Claude Opus 4.8 | Anthropic | S | $5 / $25 | 1M |
| 4 | Claude Opus 4.6 | Anthropic | S | $5 / $25 | 1M |
| 5 | Claude Opus 4.5 | Anthropic | S | $5 / $25 | 200K |
| 6 | DeepSeek V4-Pro | DeepSeek | S | $0.435 / $0.87 | 1M |
| 7 | Claude Opus 5 | Anthropic | S | $5 / $25 | 1M |
| 8 | Gemini 3.8 Flash | S | $0.75 / $3.75 | 1M | |
| 9 | Kimi K2.5 | Moonshot AI | S | $0.6 / $3 | 262K |
| 10 | Claude Opus 5.5 | Anthropic | S | $4 / $20 | 1M |
Grades run from S (best) to D. Prices are in US dollars per one million tokens, input and output.
Coding
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Anthropic | S | $4 / $20 | 1M |
| 2 | GPT-6 Astra | OpenAI | S | $10 / $50 | 1M |
| 3 | Claude Fable 5 | Anthropic | S | $10 / $50 | 1M |
| 4 | DeepSeek V4-Flash | DeepSeek | S | $0.14 / $0.28 | 1M |
| 5 | Claude Fable 5.1 | Anthropic | S | $10 / $50 | 1M |
Coding grades lean on SWE-bench Verified, which tests fixing real GitHub issues, on Arena Elo for web development, and on LiveCodeBench. See Best LLM for Coding for the full table.
Reasoning
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | S | $2 / $12 | 1M | |
| 2 | GPT-6 Astra | OpenAI | S | $10 / $50 | 1M |
| 3 | Claude Fable 5 | Anthropic | S | $10 / $50 | 1M |
| 4 | Claude Opus 4.8 | Anthropic | S | $5 / $25 | 1M |
| 5 | Grok 4.3 | xAI | S | $1.25 / $2.5 | 1M |
Reasoning grades use GPQA Diamond, Humanity's Last Exam and MMLU-Pro. The full view is at Best LLM for Reasoning.
Math
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra | OpenAI | S | $10 / $50 | 1M |
| 2 | Claude Sonnet 5.5 | Anthropic | S | $2 / $10 | 1M |
| 3 | GPT-6.1 Sol | OpenAI | S | $2 / $10 | 1.1M |
| 4 | Gemini 3.8 Flash | S | $0.75 / $3.75 | 1M | |
| 5 | GPT-5.6 Sol | OpenAI | S | $4 / $20 | 1M |
Chat and writing
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Anthropic | S | $4 / $20 | 1M |
| 2 | Claude Fable 5.1 | Anthropic | S | $10 / $50 | 1M |
| 3 | Claude Fable 5 | Anthropic | S | $10 / $50 | 1M |
| 4 | Gemini 3.8 Flash | S | $0.75 / $3.75 | 1M | |
| 5 | Gemini 3.7 Flash | S | $0.75 / $3.75 | 1M |
Chat and writing are judged mainly by Arena Elo, where people vote between two anonymous answers.
Agents and tool calling
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Anthropic | S | $4 / $20 | 1M |
| 2 | GPT-6 Astra | OpenAI | S | $10 / $50 | 1M |
| 3 | Claude Fable 5 | Anthropic | S | $10 / $50 | 1M |
| 4 | Claude Opus 5 | Anthropic | S | $5 / $25 | 1M |
| 5 | Claude Opus 4.8 | Anthropic | S | $5 / $25 | 1M |
Best open-weight models
Open-weight models can be downloaded and run on your own hardware or through any hosting provider.
| # | Model | Provider | Grade | Price in / out per 1M | Context |
|---|---|---|---|---|---|
| 1 | DeepSeek V4-Pro | DeepSeek | S | $0.435 / $0.87 | 1M |
| 2 | Kimi K2.5 | Moonshot AI | S | $0.6 / $3 | 262K |
| 3 | Kimi K3 | Moonshot AI | S | $3 / $15 | 1M |
| 4 | GLM-5.1 | Zhipu AI | S | $1.4 / $4.4 | 203K |
| 5 | Kimi K2.6 | Moonshot AI | S | $0.95 / $4 | 256K |
| 6 | MiMo-V2.5-Pro | Xiaomi | S | $0.435 / $0.87 | 1M |
| 7 | MiMo-V2.6-Flash | Xiaomi | S | $0.14 / $0.28 | 1M |
| 8 | MiMo-V2.5 | Xiaomi | S | $0.14 / $0.28 | 1M |
To see whether one fits your GPU, open it in Can I Run LLM?.
Best value: cheapest models graded A or better
| Model | Provider | Input per 1M | Output per 1M | Grade | Context |
|---|---|---|---|---|---|
| Qwen3.7 Flash | Alibaba | $0.03 | $0.13 | A | 1M |
| Qwen3.5 9B | Alibaba | $0.1 | $0.15 | A | 262K |
| DeepSeek V4-Flash | DeepSeek | $0.14 | $0.28 | A | 1M |
| MiMo-V2.5 | Xiaomi | $0.14 | $0.28 | S | 1M |
| MiMo-V2.6-Flash | Xiaomi | $0.14 | $0.28 | S | 1M |
| GPT-6 Luna | OpenAI | $0.1 | $0.5 | A | 1.1M |
| GLM-5.3 Flash | Zhipu AI | $0.15 | $0.5 | A | 1M |
| DeepSeek V3.2 | DeepSeek | $0.28 | $0.42 | A | 164K |
A cheap model that grades A is often the right choice for high-volume work. The flagship models earn their price on the hardest tasks, not on everyday ones.
How to read these tables
- A grade is a summary, not a score. Two models with the same grade are close enough that price, speed and context window should decide.
- Missing benchmark scores are common for new models. The ranking estimates a missing score from the model's other scores instead of treating it as zero.
- Prices change. Each price here was confirmed by two independent sources on the date shown at the top of this page.
Go deeper
- Compare any two models side by side.
- See the method on the Data and Methodology page.
- Read what the benchmarks measure in LLM Benchmark Scores Explained.
Frequently asked questions
Which LLM is best overall right now?
As of October 1, 2026, the top three overall on the id8 leaderboard are Claude Fable 5, Gemini 3.1 Pro and Claude Opus 4.8.
Which open-weight LLM is best right now?
The highest-ranked open-weight models overall are DeepSeek V4-Pro, Kimi K2.5 and Kimi K3.
How often does this page change?
The tables are rebuilt whenever the underlying data changes. Prices, context windows and scores are re-checked daily, and a value is published only when two independent sources agree.
How are models ranked?
First by a task grade from S to D, then by the average of the benchmarks that matter for that task. The method is described on the Data and Methodology page.