LLM Comparisons
Head-to-Head LLM Comparisons
The id8 LLM Compare hub is a free collection of 10 head-to-head AI model comparisons — benchmark scores, per-task grades, pricing, and a clear winner recommendation. Every score links to its primary source so you can verify the claim yourself.
How Each Comparison Is Built
Every comparison page covers the same five dimensions so you can evaluate consistently across pairs. Benchmark scores are pulled from Papers With Code, LMSYS Chatbot Arena, and official model cards — with a direct link to each source. Pricing is verified from provider pricing pages at the time of publication.
Benchmarks
Up to 8 scored: MMLU-Pro, SWE-bench, GPQA, HumanEval, Arena Elo
Task Grades
A–D per task type: coding, reasoning, writing, math, multilingual
Pricing
Input/output cost per million tokens (USD) from official pages
Local Run
VRAM needed per quantization level; whether self-host is viable
Verdict
Use-case-specific winner recommendation, not a single winner
Latest Model Comparisons
Head-to-head pages for the models currently at the top of the leaderboard. Prices and scores come from the same data, re-checked daily.
Gemma 4 31B vs Qwen3.8-27B
Gemma 4 31B (open weights, self-hosted) against Qwen3.8-27B (open weights, self-hosted): benchmarks, context window and task grades side by side.
DeepSeek V4.1 Flash vs GLM-5.3 Flash
DeepSeek V4.1 Flash (open weights, self-hosted) against GLM-5.3 Flash (open weights, self-hosted): benchmarks, context window and task grades side by side.
DeepSeek V4-Pro vs Kimi K3
DeepSeek V4-Pro (open weights, self-hosted) against Kimi K3 (open weights, self-hosted): benchmarks, context window and task grades side by side.
Qwen3.8 Max vs Kimi K3
Qwen3.8 Max (open weights, self-hosted) against Kimi K3 (open weights, self-hosted): benchmarks, context window and task grades side by side.
Grok 4.7 vs GPT-6 Sol
Grok 4.7 (open weights, self-hosted) against GPT-6 Sol (open weights, self-hosted): benchmarks, context window and task grades side by side.
Gemini 3.8 Flash vs GPT-6 Luna
Gemini 3.8 Flash (open weights, self-hosted) against GPT-6 Luna (open weights, self-hosted): benchmarks, context window and task grades side by side.
Claude Sonnet 5.5 vs GPT-6.1 Sol
Claude Sonnet 5.5 (open weights, self-hosted) against GPT-6.1 Sol (open weights, self-hosted): benchmarks, context window and task grades side by side.
Claude Fable 5.1 vs GPT-6 Astra
Claude Fable 5.1 (open weights, self-hosted) against GPT-6 Astra (open weights, self-hosted): benchmarks, context window and task grades side by side.
Claude Fable 5.1 vs Claude Opus 5.5
Claude Fable 5.1 (open weights, self-hosted) against Claude Opus 5.5 (open weights, self-hosted): benchmarks, context window and task grades side by side.
Claude Opus 5.5 vs GPT-6 Astra
Claude Opus 5.5 (open weights, self-hosted) against GPT-6 Astra (open weights, self-hosted): benchmarks, context window and task grades side by side.
MiniMax M3 vs Qwen 3.5
MiniMax M3 (open weights, self-hosted) against Qwen 3.5 (open weights, self-hosted): benchmarks, context window and task grades side by side.
GLM-5.2 vs MiniMax M3
GLM-5.2 (open weights, self-hosted) against MiniMax M3 (open weights, self-hosted): benchmarks, context window and task grades side by side.
Claude Sonnet 5 vs GPT-5.5 Instant
Claude Sonnet 5 (open weights, self-hosted) against GPT-5.5 Instant (open weights, self-hosted): benchmarks, context window and task grades side by side.
Flagship API Model Comparisons
The top-tier paid models from Anthropic, OpenAI, and Google — for teams evaluating which provider to build on.
Claude Opus 4.7 vs GPT-5.5 Instant
Premium reasoning power vs rapid response speed — the flagship API showdown for enterprise use.
Gemini 3.1 Flash Lite vs GPT-5.5 Instant
Ultra-fast, cost-efficient API models — which gives more for less at high throughput?
Claude Sonnet 4.6 vs DeepSeek V3.2
Mid-tier API model vs the best open-source challenger. Benchmark scores, pricing, and coding tasks compared.
Open-Source Model Comparisons
Self-hostable models you can run on your own GPU — compared on benchmarks, VRAM requirements, and task performance.
DeepSeek R1 (Full) vs Nemotron 3 Ultra 550B
Open-source reasoning giants: 671B vs 550B parameters. Which self-hosted heavyweight wins on MMLU-Pro and SWE-bench?
Llama 4 Maverick vs DeepSeek R1 (Full)
Meta's MoE flagship vs DeepSeek's best — the defining open-source matchup of 2026.
Gemma 4 26B (MoE) vs Qwen 3 32B
Google vs Alibaba in the mid-size open-source tier — fits on a single RTX 4090, but which runs better tasks?
Llama 3.3 70B vs Llama 4 Scout
Meta's generations head-to-head: is the Llama 4 upgrade worth switching for your use case?
Coding Specialist Comparisons
Models fine-tuned or optimised specifically for code generation — ranked by SWE-bench Verified, HumanEval, and LiveCodeBench.
Compact & Efficient Model Comparisons
Sub-30B models that run on a single consumer GPU — compared for teams where VRAM budget or inference cost is the constraint.
Frequently Asked Questions
Which LLM benchmarks do the comparisons use?
Each comparison uses up to 8 benchmarks depending on model availability: MMLU-Pro (expert-level knowledge), SWE-bench Verified (real-world coding), GPQA Diamond (PhD-level reasoning), HumanEval (function synthesis), LiveCodeBench (competitive programming), Chatbot Arena Elo (human preference), MATH-500 (competition mathematics), and GSM8K (word problems). Every score links to its primary source.
How is the winner determined in each comparison?
There is no single winner. Each page assigns per-task grades (A/B/C/D) across coding, reasoning, writing, math, and cost-efficiency. The verdict section picks the better model for specific use cases — one model may win on coding while another wins on price-per-token.
Are the benchmark scores up to date?
Comparison pages are updated when major benchmark results change. Each page shows a last-modified date. For the most current scores across 93 models use the id8 LLM Leaderboard — updated monthly.
Can I compare models not listed here?
The id8 LLM Leaderboard ranks 93 models across 20 benchmarks with filter-by-task support. You can filter to any two models and compare their scores side by side. Dedicated comparison pages for additional pairs are added regularly.
What is the difference between API models and open-source models?
API models (Claude, GPT, Gemini) are accessed through a paid API — you pay per token and the model runs on the provider's servers. Open-source models (Llama, DeepSeek, Gemma, Qwen) can be downloaded and run locally on your own GPU. The comparisons note which models are self-hostable and include VRAM requirements. Use the VRAM Calculator to check if your GPU can run a specific model.
Can't find your pair?
The full leaderboard ranks 93 models across 20 benchmarks — filter by task, budget, or self-host availability to find your winner.