Pricing Guide

Cheapest LLM API in 2026: Price per Million Tokens, Ranked

Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset

The cheapest model is rarely the best choice and the best model is rarely needed. This page lists the lowest-priced APIs that still earn an A or S grade on the id8 leaderboard, as of October 2, 2026, and explains how to estimate what you will really pay.

Input, output and average price for all models, with a budget filter.

Sort every model by price →

Cheapest models graded A or better, overall

ModelProviderInput per 1MOutput per 1MGradeContext
Qwen3.7 FlashAlibaba$0.03$0.13A1M
Qwen3.5 9BAlibaba$0.1$0.15A262K
DeepSeek V4-FlashDeepSeek$0.14$0.28A1M
MiMo-V2.5Xiaomi$0.14$0.28S1M
MiMo-V2.6-FlashXiaomi$0.14$0.28S1M
GPT-6 LunaOpenAI$0.1$0.5A1.1M
GLM-5.3 FlashZhipu AI$0.15$0.5A1M
DeepSeek V3.2DeepSeek$0.28$0.42A164K
GPT-oss 120BOpenAI$0.15$0.6A131K
DeepSeek V4.1 FlashDeepSeek$0.15$0.6A1M

Prices are US dollars per one million tokens. Only models with an overall grade of A or S are listed, ordered by input plus output price.

Cheapest for coding

ModelProviderInput per 1MOutput per 1MGradeContext
Qwen3.7 FlashAlibaba$0.03$0.13A1M
DeepSeek V4-FlashDeepSeek$0.14$0.28S1M
MiMo-V2.5Xiaomi$0.14$0.28S1M
MiMo-V2.6-FlashXiaomi$0.14$0.28S1M
GPT-6 LunaOpenAI$0.1$0.5S1.1M
GLM-5.3 FlashZhipu AI$0.15$0.5A1M
DeepSeek V3.2DeepSeek$0.28$0.42A164K
GPT-oss 120BOpenAI$0.15$0.6A131K

Cheapest for chat

ModelProviderInput per 1MOutput per 1MGradeContext
Qwen3.7 FlashAlibaba$0.03$0.13A1M
DeepSeek V4-FlashDeepSeek$0.14$0.28A1M
MiMo-V2.5Xiaomi$0.14$0.28A1M
MiMo-V2.6-FlashXiaomi$0.14$0.28S1M
GPT-6 LunaOpenAI$0.1$0.5A1.1M
GPT-oss 120BOpenAI$0.15$0.6A131K
DeepSeek V4.1 FlashDeepSeek$0.15$0.6A1M
DeepSeek V4-ProDeepSeek$0.435$0.87A1M

Cheapest for agents and tool use

ModelProviderInput per 1MOutput per 1MGradeContext
Qwen3.7 FlashAlibaba$0.03$0.13A1M
DeepSeek V4-FlashDeepSeek$0.14$0.28A1M
MiMo-V2.5Xiaomi$0.14$0.28A1M
MiMo-V2.6-FlashXiaomi$0.14$0.28A1M
GPT-6 LunaOpenAI$0.1$0.5A1.1M
GLM-5.3 FlashZhipu AI$0.15$0.5A1M
DeepSeek V4.1 FlashDeepSeek$0.15$0.6S1M
DeepSeek V4-ProDeepSeek$0.435$0.87S1M

Agents make many calls per task, so price per token matters more for agents than for any other use.

How to work out what you will pay

A token is roughly three quarters of an English word. Hindi and other Indic scripts use more tokens per word.

Cost per request = (input tokens × input price + output tokens × output price) ÷ 1,000,000

Worked example: a support bot sends 1,500 tokens of instructions and history and receives 300 tokens back. With a model priced at $0.50 input and $2.00 output per million tokens:

  • Input: 1,500 × 0.50 ÷ 1,000,000 = $0.00075
  • Output: 300 × 2.00 ÷ 1,000,000 = $0.00060
  • Total: $0.00135 per request, or $1.35 per thousand requests

Run the same sum with the prices in the tables above for the models you are considering.

What changes the bill more than the model

  • Prompt length. A long system prompt is paid for on every request. Trim it. Prompt Zipper removes filler words from prompts.
  • Conversation history. Sending the whole chat every turn grows cost with each message. Summarise older turns.
  • Prompt caching. Most providers discount input tokens that repeat at the start of a request. Put fixed instructions first to benefit.
  • Batch processing. Jobs that can wait are often billed at a discount.
  • Output limits. Set a maximum output length so a runaway answer cannot cost ten times the usual.
  • Reasoning modes. Thinking tokens are billed as output. Switch reasoning on only for tasks that need it.

Route by difficulty

The most effective saving is not choosing one cheap model; it is using two:

  1. Send every request to a cheap model graded A.
  2. Send only the requests it handles badly, judged by a rule or by a quick check, to a flagship.

For most products the cheap model handles the large majority of requests.

Open-weight models through hosting providers

Open-weight models are served by many providers, who compete on price. The strongest open-weight models overall right now are DeepSeek V4-Pro, Kimi K2.5 and Kimi K3. If you later need to self-host for privacy, the same model can run on your own hardware; check requirements in Can I Run LLM?.

A note on these prices

Each price is published only when two independent sources agree, and is re-checked daily. Promotional discounts and volume pricing are not included. Confirm on the provider's pricing page before you commit.

Frequently asked questions

What is the cheapest good LLM API right now?

The cheapest model graded A or better overall is Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens), as of October 2, 2026.

What is the cheapest LLM API for coding?

The cheapest model graded A or better for coding is Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens).

Why are output tokens more expensive than input tokens?

Generating text takes one model pass per token, while input tokens are processed together. Providers price output several times higher to reflect that.

Is a free tier enough for a real product?

Free tiers have low rate limits and may use your data for training. They are fine for prototypes; a product needs a paid tier with clear data terms.

Is running a model myself cheaper?

Only at high, steady volume. A rented GPU costs the same whether it is busy or idle, so self-hosting pays off when the hardware is used most of the day.

Related