LLM Comparisons

Head-to-Head LLM Comparisons

The id8 LLM Compare hub is a free collection of 10 head-to-head AI model comparisons — benchmark scores, per-task grades, pricing, and a clear winner recommendation. Every score links to its primary source so you can verify the claim yourself.

How Each Comparison Is Built

Every comparison page covers the same five dimensions so you can evaluate consistently across pairs. Benchmark scores are pulled from Papers With Code, LMSYS Chatbot Arena, and official model cards — with a direct link to each source. Pricing is verified from provider pricing pages at the time of publication.

Benchmarks

Up to 8 scored: MMLU-Pro, SWE-bench, GPQA, HumanEval, Arena Elo

Task Grades

A–D per task type: coding, reasoning, writing, math, multilingual

Pricing

Input/output cost per million tokens (USD) from official pages

Local Run

VRAM needed per quantization level; whether self-host is viable

Verdict

Use-case-specific winner recommendation, not a single winner

Latest Model Comparisons

Head-to-head pages for the models currently at the top of the leaderboard. Prices and scores come from the same data, re-checked daily.

Google vs Alibaba

Gemma 4 31B vs Qwen3.8-27B

Gemma 4 31B (open weights, self-hosted) against Qwen3.8-27B (open weights, self-hosted): benchmarks, context window and task grades side by side.

DeepSeek vs Zhipu AI

DeepSeek V4.1 Flash vs GLM-5.3 Flash

DeepSeek V4.1 Flash (open weights, self-hosted) against GLM-5.3 Flash (open weights, self-hosted): benchmarks, context window and task grades side by side.

DeepSeek vs Moonshot AI

DeepSeek V4-Pro vs Kimi K3

DeepSeek V4-Pro (open weights, self-hosted) against Kimi K3 (open weights, self-hosted): benchmarks, context window and task grades side by side.

Alibaba vs Moonshot AI

Qwen3.8 Max vs Kimi K3

Qwen3.8 Max (open weights, self-hosted) against Kimi K3 (open weights, self-hosted): benchmarks, context window and task grades side by side.

xAI vs OpenAI

Grok 4.7 vs GPT-6 Sol

Grok 4.7 (open weights, self-hosted) against GPT-6 Sol (open weights, self-hosted): benchmarks, context window and task grades side by side.

Google vs OpenAI

Gemini 3.8 Flash vs GPT-6 Luna

Gemini 3.8 Flash (open weights, self-hosted) against GPT-6 Luna (open weights, self-hosted): benchmarks, context window and task grades side by side.

Anthropic vs OpenAI

Claude Sonnet 5.5 vs GPT-6.1 Sol

Claude Sonnet 5.5 (open weights, self-hosted) against GPT-6.1 Sol (open weights, self-hosted): benchmarks, context window and task grades side by side.

Anthropic vs OpenAI

Claude Fable 5.1 vs GPT-6 Astra

Claude Fable 5.1 (open weights, self-hosted) against GPT-6 Astra (open weights, self-hosted): benchmarks, context window and task grades side by side.

Anthropic vs Anthropic

Claude Fable 5.1 vs Claude Opus 5.5

Claude Fable 5.1 (open weights, self-hosted) against Claude Opus 5.5 (open weights, self-hosted): benchmarks, context window and task grades side by side.

Anthropic vs OpenAI

Claude Opus 5.5 vs GPT-6 Astra

Claude Opus 5.5 (open weights, self-hosted) against GPT-6 Astra (open weights, self-hosted): benchmarks, context window and task grades side by side.

MiniMax vs Alibaba

MiniMax M3 vs Qwen 3.5

MiniMax M3 (open weights, self-hosted) against Qwen 3.5 (open weights, self-hosted): benchmarks, context window and task grades side by side.

Zhipu AI vs MiniMax

GLM-5.2 vs MiniMax M3

GLM-5.2 (open weights, self-hosted) against MiniMax M3 (open weights, self-hosted): benchmarks, context window and task grades side by side.

Anthropic vs OpenAI

Claude Sonnet 5 vs GPT-5.5 Instant

Claude Sonnet 5 (open weights, self-hosted) against GPT-5.5 Instant (open weights, self-hosted): benchmarks, context window and task grades side by side.

Flagship API Model Comparisons

The top-tier paid models from Anthropic, OpenAI, and Google — for teams evaluating which provider to build on.

Open-Source Model Comparisons

Self-hostable models you can run on your own GPU — compared on benchmarks, VRAM requirements, and task performance.

Coding Specialist Comparisons

Models fine-tuned or optimised specifically for code generation — ranked by SWE-bench Verified, HumanEval, and LiveCodeBench.

Compact & Efficient Model Comparisons

Sub-30B models that run on a single consumer GPU — compared for teams where VRAM budget or inference cost is the constraint.

Frequently Asked Questions

Which LLM benchmarks do the comparisons use?

Each comparison uses up to 8 benchmarks depending on model availability: MMLU-Pro (expert-level knowledge), SWE-bench Verified (real-world coding), GPQA Diamond (PhD-level reasoning), HumanEval (function synthesis), LiveCodeBench (competitive programming), Chatbot Arena Elo (human preference), MATH-500 (competition mathematics), and GSM8K (word problems). Every score links to its primary source.

How is the winner determined in each comparison?

There is no single winner. Each page assigns per-task grades (A/B/C/D) across coding, reasoning, writing, math, and cost-efficiency. The verdict section picks the better model for specific use cases — one model may win on coding while another wins on price-per-token.

Are the benchmark scores up to date?

Comparison pages are updated when major benchmark results change. Each page shows a last-modified date. For the most current scores across 93 models use the id8 LLM Leaderboard — updated monthly.

Can I compare models not listed here?

The id8 LLM Leaderboard ranks 93 models across 20 benchmarks with filter-by-task support. You can filter to any two models and compare their scores side by side. Dedicated comparison pages for additional pairs are added regularly.

What is the difference between API models and open-source models?

API models (Claude, GPT, Gemini) are accessed through a paid API — you pay per token and the model runs on the provider's servers. Open-source models (Llama, DeepSeek, Gemma, Qwen) can be downloaded and run locally on your own GPU. The comparisons note which models are self-hostable and include VRAM requirements. Use the VRAM Calculator to check if your GPU can run a specific model.

Can't find your pair?

The full leaderboard ranks 93 models across 20 benchmarks — filter by task, budget, or self-host availability to find your winner.

Open Leaderboard