Cheapest LLM APIs by Output Price

79 models qualify · sorted by output price · verified 2 October 2026

The cheapest is ERNIE X1 at $0.02 per 1M output tokens (Baidu ERNIE (百度)). Sorted, current, and computed from the published dataset rather than hand-maintained.

Before you switch. This is the column that decides most bills. If a model is cheap to read from and expensive to write to, a chat workload will cost you far more than the input rate suggests.

Full ranking

Top 15 of 79, ascending by output price per 1M tokens. Verified 2 October 2026.
ModelProviderCountingContext Input / 1MOutput / 1MMax output
ERNIE X1Baidu ERNIE (百度)estimate128K$0.006$0.02—
Doubao 1.6 LiteByteDance Doubao (火山引擎)estimate128K$0.02$0.04—
Doubao 1.6 FlashByteDance Doubao (火山引擎)estimate128K$0.01$0.11—
DeepSeek V4 FlashDeepSeek (深度求索)estimate400K$0.14$0.28—
MiMo V2.5Xiaomi MiMo (小米)estimate1M$0.14$0.28—
GPT-5 nanoOpenAIexact400K$0.05$0.4—
GPT-4.1 nanoOpenAIexact1M$0.1$0.4—
Gemini 2.5 Flash-LiteGoogle Geminiestimate1M$0.1$0.4—
DeepSeek V3.2DeepSeek (深度求索)estimate128K$0.35$0.42—
Llama 4 (Meta)US & EUestimate1M$0.15$0.45—
Doubao 1.6 Pro (豆包)ByteDance Doubao (火山引擎)estimate256K$0.11$0.45—
ERNIE 4.5 (文心一言)Baidu ERNIE (百度)estimate128K$0.2$0.55—
GPT-4o miniOpenAIexact128K$0.15$0.6—
SenseNova 5.1 (商汤)SenseTime (商汤)estimate128K$0.04$0.7—
DeepSeek V4 ProDeepSeek (深度求索)estimate128K$0.44$0.87—

What a request actually costs

At $0.02 per 1M output tokens, ERNIE X1 costs:

  • 1M tokens — $0.02
  • 100M tokens — $2.00
  • 1B tokens — $20.00

Cheap per token stops being cheap once volume compounds. At a billion tokens the difference between the top two rows on this page is tens of thousands of dollars a year.

Prefer to check one model in detail? All 84 model calculators, every provider's pricing, or the raw JSON dataset.

FAQs

What is the cheapest output per 1M tokens right now?

ERNIE X1 at $0.02 per 1M output tokens, from Baidu ERNIE (百度), as verified on 2 October 2026.

Are these list prices?

Yes. Public list prices only. Batch rates, committed-use discounts and enterprise contracts are cheaper and are not modelled here. Check the verified date on each page before you budget against it.

Are the token counts exact?

Only for OpenAI models, which are counted with the same tiktoken tokenizer the API bills with. Every other vendor publishes no tokenizer, so those counts are calibrated estimates at 5-10% accuracy.

Does a cheap input price make a model cheap overall?

Not necessarily. On this page the ranking is by output price. Output rates run 3-8x input rates on most models, so a chat workload is usually an output workload.