Monthly Update

LLM Leaderboard Update, October 2026: Who Leads Each Task Now

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

A plain summary of where the id8 LLM Leaderboard stands as of October 1, 2026: the leaders overall and by task, the strongest open-weight models, and the cheapest models that still grade A or better. The tables on this page are generated from the same data as the leaderboard.

All models, every benchmark, prices and context windows. Free, no signup.

Open the full leaderboard →

The short version

  • Overall: Claude Fable 5, Gemini 3.1 Pro and Claude Opus 4.8 lead.
  • Coding: Claude Opus 5.5, GPT-6 Astra and Claude Fable 5.
  • Reasoning: Gemini 3.1 Pro, GPT-6 Astra and Claude Fable 5.
  • Best open-weight overall: DeepSeek V4-Pro, Kimi K2.5 and Kimi K3.
  • Cheapest model graded A or better overall: Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens).

The leaderboard currently tracks 164 models, of which 95 can be run on your own hardware.

Top 10 overall

#ModelProviderGradePrice in / out per 1MContext
1Claude Fable 5AnthropicS$10 / $501M
2Gemini 3.1 ProGoogleS$2 / $121M
3Claude Opus 4.8AnthropicS$5 / $251M
4Claude Opus 4.6AnthropicS$5 / $251M
5Claude Opus 4.5AnthropicS$5 / $25200K
6DeepSeek V4-ProDeepSeekS$0.435 / $0.871M
7Claude Opus 5AnthropicS$5 / $251M
8Gemini 3.8 FlashGoogleS$0.75 / $3.751M
9Kimi K2.5Moonshot AIS$0.6 / $3262K
10Claude Opus 5.5AnthropicS$4 / $201M

Grades run from S (best) to D. Prices are in US dollars per one million tokens, input and output.

Coding

#ModelProviderGradePrice in / out per 1MContext
1Claude Opus 5.5AnthropicS$4 / $201M
2GPT-6 AstraOpenAIS$10 / $501M
3Claude Fable 5AnthropicS$10 / $501M
4DeepSeek V4-FlashDeepSeekS$0.14 / $0.281M
5Claude Fable 5.1AnthropicS$10 / $501M

Coding grades lean on SWE-bench Verified, which tests fixing real GitHub issues, on Arena Elo for web development, and on LiveCodeBench. See Best LLM for Coding for the full table.

Reasoning

#ModelProviderGradePrice in / out per 1MContext
1Gemini 3.1 ProGoogleS$2 / $121M
2GPT-6 AstraOpenAIS$10 / $501M
3Claude Fable 5AnthropicS$10 / $501M
4Claude Opus 4.8AnthropicS$5 / $251M
5Grok 4.3xAIS$1.25 / $2.51M

Reasoning grades use GPQA Diamond, Humanity's Last Exam and MMLU-Pro. The full view is at Best LLM for Reasoning.

Math

#ModelProviderGradePrice in / out per 1MContext
1GPT-6 AstraOpenAIS$10 / $501M
2Claude Sonnet 5.5AnthropicS$2 / $101M
3GPT-6.1 SolOpenAIS$2 / $101.1M
4Gemini 3.8 FlashGoogleS$0.75 / $3.751M
5GPT-5.6 SolOpenAIS$4 / $201M

Chat and writing

#ModelProviderGradePrice in / out per 1MContext
1Claude Opus 5.5AnthropicS$4 / $201M
2Claude Fable 5.1AnthropicS$10 / $501M
3Claude Fable 5AnthropicS$10 / $501M
4Gemini 3.8 FlashGoogleS$0.75 / $3.751M
5Gemini 3.7 FlashGoogleS$0.75 / $3.751M

Chat and writing are judged mainly by Arena Elo, where people vote between two anonymous answers.

Agents and tool calling

#ModelProviderGradePrice in / out per 1MContext
1Claude Opus 5.5AnthropicS$4 / $201M
2GPT-6 AstraOpenAIS$10 / $501M
3Claude Fable 5AnthropicS$10 / $501M
4Claude Opus 5AnthropicS$5 / $251M
5Claude Opus 4.8AnthropicS$5 / $251M

Best open-weight models

Open-weight models can be downloaded and run on your own hardware or through any hosting provider.

#ModelProviderGradePrice in / out per 1MContext
1DeepSeek V4-ProDeepSeekS$0.435 / $0.871M
2Kimi K2.5Moonshot AIS$0.6 / $3262K
3Kimi K3Moonshot AIS$3 / $151M
4GLM-5.1Zhipu AIS$1.4 / $4.4203K
5Kimi K2.6Moonshot AIS$0.95 / $4256K
6MiMo-V2.5-ProXiaomiS$0.435 / $0.871M
7MiMo-V2.6-FlashXiaomiS$0.14 / $0.281M
8MiMo-V2.5XiaomiS$0.14 / $0.281M

To see whether one fits your GPU, open it in Can I Run LLM?.

Best value: cheapest models graded A or better

ModelProviderInput per 1MOutput per 1MGradeContext
Qwen3.7 FlashAlibaba$0.03$0.13A1M
Qwen3.5 9BAlibaba$0.1$0.15A262K
DeepSeek V4-FlashDeepSeek$0.14$0.28A1M
MiMo-V2.5Xiaomi$0.14$0.28S1M
MiMo-V2.6-FlashXiaomi$0.14$0.28S1M
GPT-6 LunaOpenAI$0.1$0.5A1.1M
GLM-5.3 FlashZhipu AI$0.15$0.5A1M
DeepSeek V3.2DeepSeek$0.28$0.42A164K

A cheap model that grades A is often the right choice for high-volume work. The flagship models earn their price on the hardest tasks, not on everyday ones.

How to read these tables

  • A grade is a summary, not a score. Two models with the same grade are close enough that price, speed and context window should decide.
  • Missing benchmark scores are common for new models. The ranking estimates a missing score from the model's other scores instead of treating it as zero.
  • Prices change. Each price here was confirmed by two independent sources on the date shown at the top of this page.

Go deeper

Frequently asked questions

Which LLM is best overall right now?

As of October 1, 2026, the top three overall on the id8 leaderboard are Claude Fable 5, Gemini 3.1 Pro and Claude Opus 4.8.

Which open-weight LLM is best right now?

The highest-ranked open-weight models overall are DeepSeek V4-Pro, Kimi K2.5 and Kimi K3.

How often does this page change?

The tables are rebuilt whenever the underlying data changes. Prices, context windows and scores are re-checked daily, and a value is published only when two independent sources agree.

How are models ranked?

First by a task grade from S to D, then by the average of the benchmarks that matter for that task. The method is described on the Data and Methodology page.

Related