Cheapest Models With 1M+ Token Context
The cheapest is MiMo V2.5 at $0.28 per 1M output tokens (Xiaomi MiMo (小米)). Sorted, current, and computed from the published dataset rather than hand-maintained.
Full ranking
| Model | Provider | Counting | Context | Input / 1M | Output / 1M | Max output |
|---|---|---|---|---|---|---|
| MiMo V2.5 | Xiaomi MiMo (小米) | estimate | 1M | $0.14 | $0.28 | — |
| GPT-4.1 nano | OpenAI | exact | 1M | $0.1 | $0.4 | — |
| Gemini 2.5 Flash-Lite | Google Gemini | estimate | 1M | $0.1 | $0.4 | — |
| Llama 4 (Meta) | US & EU | estimate | 1M | $0.15 | $0.45 | — |
| MiMo V2.5 Pro (小米) | Xiaomi MiMo (小米) | estimate | 1M | $0.44 | $0.87 | — |
| Grok 4.1 (xAI) | US & EU | estimate | 1M | $0.2 | $1.00 | — |
| GPT-5.6 Luna | OpenAI | exact | 1M | $0.2 | $1.20 | — |
| SkyClaw 1.0 (昆仑万维) | Kunlun (昆仑万维) | estimate | 1M | $0.5 | $1.50 | — |
| GPT-4.1 mini | OpenAI | exact | 1M | $0.4 | $1.60 | — |
| Gemini 3.5 Flash-Lite | Google Gemini | estimate | 1M | $0.3 | $2.50 | — |
| Gemini 2.5 Flash | Google Gemini | estimate | 1M | $0.3 | $2.50 | — |
| Grok 4.3 | US & EU | estimate | 1M | $1.25 | $2.50 | 32K |
| Gemini 3 Flash | Google Gemini | estimate | 1M | $0.5 | $3.00 | — |
| Gemini 3.6 Flash | Google Gemini | estimate | 1M | $1.50 | $7.50 | — |
| GPT-4.1 | OpenAI | exact | 1M | $2.00 | $8.00 | — |
What a request actually costs
At $0.28 per 1M output tokens, MiMo V2.5 costs:
- 1M tokens — $0.28
- 100M tokens — $28.00
- 1B tokens — $280.00
Cheap per token stops being cheap once volume compounds. At a billion tokens the difference between the top two rows on this page is tens of thousands of dollars a year.
Prefer to check one model in detail? All 84 model calculators, every provider's pricing, or the raw JSON dataset.
FAQs
What is the cheapest output per 1M tokens right now?
MiMo V2.5 at $0.28 per 1M output tokens, from Xiaomi MiMo (小米), as verified on 2 October 2026.
Are these list prices?
Yes. Public list prices only. Batch rates, committed-use discounts and enterprise contracts are cheaper and are not modelled here. Check the verified date on each page before you budget against it.
Are the token counts exact?
Only for OpenAI models, which are counted with the same tiktoken tokenizer the API bills with. Every other vendor publishes no tokenizer, so those counts are calibrated estimates at 5-10% accuracy.
Does a cheap input price make a model cheap overall?
Not necessarily. On this page the ranking is by output price. Output rates run 3-8x input rates on most models, so a chat workload is usually an output workload.