Cheapest LLM API Pricing, Answered
Everyone else publishes a pricing table. This is a sorted answer, cut by the constraint that actually decides which model you should pick. Every row is computed from the published dataset, not typed in.
Pick your constraint
- Cheapest per 1M input tokens — long prompts, big documents, RAG corpora
- Cheapest per 1M output tokens — long generations, and this is where bills actually go
- Cheapest with 200K+ context — when the context window is the binding limit
- Cheapest with 1M+ context — whole-repo analysis, long agent histories
- Cheapest OpenAI model — and where those counts are exact rather than estimated
The number that matters is usually output. On most
models the output rate is 3–8× the input rate. A workload with short prompts
and long answers is an output workload. Rank by output unless your
prompts are the long part.