Pricing Guide
Cheapest LLM API in 2026: Price per Million Tokens, Ranked
Last updated: October 2, 2026 — model data is refreshed automatically from the id8 dataset
The cheapest model is rarely the best choice and the best model is rarely needed. This page lists the lowest-priced APIs that still earn an A or S grade on the id8 leaderboard, as of October 2, 2026, and explains how to estimate what you will really pay.
Input, output and average price for all models, with a budget filter.
Sort every model by price →Cheapest models graded A or better, overall
| Model | Provider | Input per 1M | Output per 1M | Grade | Context |
|---|---|---|---|---|---|
| Qwen3.7 Flash | Alibaba | $0.03 | $0.13 | A | 1M |
| Qwen3.5 9B | Alibaba | $0.1 | $0.15 | A | 262K |
| DeepSeek V4-Flash | DeepSeek | $0.14 | $0.28 | A | 1M |
| MiMo-V2.5 | Xiaomi | $0.14 | $0.28 | S | 1M |
| MiMo-V2.6-Flash | Xiaomi | $0.14 | $0.28 | S | 1M |
| GPT-6 Luna | OpenAI | $0.1 | $0.5 | A | 1.1M |
| GLM-5.3 Flash | Zhipu AI | $0.15 | $0.5 | A | 1M |
| DeepSeek V3.2 | DeepSeek | $0.28 | $0.42 | A | 164K |
| GPT-oss 120B | OpenAI | $0.15 | $0.6 | A | 131K |
| DeepSeek V4.1 Flash | DeepSeek | $0.15 | $0.6 | A | 1M |
Prices are US dollars per one million tokens. Only models with an overall grade of A or S are listed, ordered by input plus output price.
Cheapest for coding
| Model | Provider | Input per 1M | Output per 1M | Grade | Context |
|---|---|---|---|---|---|
| Qwen3.7 Flash | Alibaba | $0.03 | $0.13 | A | 1M |
| DeepSeek V4-Flash | DeepSeek | $0.14 | $0.28 | S | 1M |
| MiMo-V2.5 | Xiaomi | $0.14 | $0.28 | S | 1M |
| MiMo-V2.6-Flash | Xiaomi | $0.14 | $0.28 | S | 1M |
| GPT-6 Luna | OpenAI | $0.1 | $0.5 | S | 1.1M |
| GLM-5.3 Flash | Zhipu AI | $0.15 | $0.5 | A | 1M |
| DeepSeek V3.2 | DeepSeek | $0.28 | $0.42 | A | 164K |
| GPT-oss 120B | OpenAI | $0.15 | $0.6 | A | 131K |
Cheapest for chat
| Model | Provider | Input per 1M | Output per 1M | Grade | Context |
|---|---|---|---|---|---|
| Qwen3.7 Flash | Alibaba | $0.03 | $0.13 | A | 1M |
| DeepSeek V4-Flash | DeepSeek | $0.14 | $0.28 | A | 1M |
| MiMo-V2.5 | Xiaomi | $0.14 | $0.28 | A | 1M |
| MiMo-V2.6-Flash | Xiaomi | $0.14 | $0.28 | S | 1M |
| GPT-6 Luna | OpenAI | $0.1 | $0.5 | A | 1.1M |
| GPT-oss 120B | OpenAI | $0.15 | $0.6 | A | 131K |
| DeepSeek V4.1 Flash | DeepSeek | $0.15 | $0.6 | A | 1M |
| DeepSeek V4-Pro | DeepSeek | $0.435 | $0.87 | A | 1M |
Cheapest for agents and tool use
| Model | Provider | Input per 1M | Output per 1M | Grade | Context |
|---|---|---|---|---|---|
| Qwen3.7 Flash | Alibaba | $0.03 | $0.13 | A | 1M |
| DeepSeek V4-Flash | DeepSeek | $0.14 | $0.28 | A | 1M |
| MiMo-V2.5 | Xiaomi | $0.14 | $0.28 | A | 1M |
| MiMo-V2.6-Flash | Xiaomi | $0.14 | $0.28 | A | 1M |
| GPT-6 Luna | OpenAI | $0.1 | $0.5 | A | 1.1M |
| GLM-5.3 Flash | Zhipu AI | $0.15 | $0.5 | A | 1M |
| DeepSeek V4.1 Flash | DeepSeek | $0.15 | $0.6 | S | 1M |
| DeepSeek V4-Pro | DeepSeek | $0.435 | $0.87 | S | 1M |
Agents make many calls per task, so price per token matters more for agents than for any other use.
How to work out what you will pay
A token is roughly three quarters of an English word. Hindi and other Indic scripts use more tokens per word.
Cost per request = (input tokens × input price + output tokens × output price) ÷ 1,000,000
Worked example: a support bot sends 1,500 tokens of instructions and history and receives 300 tokens back. With a model priced at $0.50 input and $2.00 output per million tokens:
- Input: 1,500 × 0.50 ÷ 1,000,000 = $0.00075
- Output: 300 × 2.00 ÷ 1,000,000 = $0.00060
- Total: $0.00135 per request, or $1.35 per thousand requests
Run the same sum with the prices in the tables above for the models you are considering.
What changes the bill more than the model
- Prompt length. A long system prompt is paid for on every request. Trim it. Prompt Zipper removes filler words from prompts.
- Conversation history. Sending the whole chat every turn grows cost with each message. Summarise older turns.
- Prompt caching. Most providers discount input tokens that repeat at the start of a request. Put fixed instructions first to benefit.
- Batch processing. Jobs that can wait are often billed at a discount.
- Output limits. Set a maximum output length so a runaway answer cannot cost ten times the usual.
- Reasoning modes. Thinking tokens are billed as output. Switch reasoning on only for tasks that need it.
Route by difficulty
The most effective saving is not choosing one cheap model; it is using two:
- Send every request to a cheap model graded A.
- Send only the requests it handles badly, judged by a rule or by a quick check, to a flagship.
For most products the cheap model handles the large majority of requests.
Open-weight models through hosting providers
Open-weight models are served by many providers, who compete on price. The strongest open-weight models overall right now are DeepSeek V4-Pro, Kimi K2.5 and Kimi K3. If you later need to self-host for privacy, the same model can run on your own hardware; check requirements in Can I Run LLM?.
A note on these prices
Each price is published only when two independent sources agree, and is re-checked daily. Promotional discounts and volume pricing are not included. Confirm on the provider's pricing page before you commit.
Frequently asked questions
What is the cheapest good LLM API right now?
The cheapest model graded A or better overall is Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens), as of October 2, 2026.
What is the cheapest LLM API for coding?
The cheapest model graded A or better for coding is Qwen3.7 Flash ($0.03 / $0.13 per 1M tokens).
Why are output tokens more expensive than input tokens?
Generating text takes one model pass per token, while input tokens are processed together. Providers price output several times higher to reflect that.
Is a free tier enough for a real product?
Free tiers have low rate limits and may use your data for training. They are fine for prototypes; a product needs a paid tier with clear data terms.
Is running a model myself cheaper?
Only at high, steady volume. A rented GPU costs the same whether it is busy or idle, so self-hosting pays off when the hardware is used most of the day.