Frequently asked questions
Which LLM is best for math in 2026?
Right now GPT-6 Astra, Claude Sonnet 5.5 and GPT-6.1 Sol lead (graded on AIME 2024–25, MATH-500 and GPQA Diamond plus an overall task grade). Best open-weight: Kimi K3. Best budget pick graded A or better: Qwen3.7 Flash ($0.03/$0.13 per 1M tokens).
Why are reasoning models better at math?
Reasoning (or 'thinking') models work through a hidden chain of steps before they answer, so they can check and correct themselves. On competition math this lifts scores sharply. The cost is more output tokens and longer waits.
What are AIME and MATH-500?
AIME is the American Invitational Mathematics Examination: 15 short-answer problems per year, used as a qualifier for the US Olympiad. MATH-500 is a 500-problem sample of the MATH dataset of competition problems. Both have exact answers, so scoring is unambiguous.
Which free or open-weight LLM is best at math?
Kimi K3, DeepSeek V4-Pro and DeepSeek R1 are the top open-weight models for math right now. Use the self-host filter above the table to see which of them fit on your hardware.
Can I trust an LLM with arithmetic?
Not for long calculations. Every model makes mistakes on large-number arithmetic when working unaided. For anything numerical, let the model write and run code, and treat its own mental arithmetic as a draft.