August 3, 2026Updated September 15, 20267 min read

Cheapest LLM API per million tokens

The cheapest LLM API is not one model — it depends on your input:output ratio. Here is the budget tier ranked by blended price (updated for Qwen-Flash's September 2026 price cut, which pulls a current-generation model level with the cheapest legacy one), a worked example at 10M+2M tokens/month, and the fine print that decides the real winner.

The cheapest LLM API per million tokens is not a single model — it depends on how your traffic splits between input and output — but the budget field clusters tightly at the bottom, and one legacy model still undercuts the whole current generation. This is that tier, ranked, with the fine print that decides the real winner.

To rank models that price input and output differently on one axis, we use a blended price: (3 × input + 1 × output) ÷ 4, a rough 3-to-1 input-to-output mix typical of RAG and chat. Your real ratio changes the order — an output-heavy workload punishes high output rates far more — so treat the ranking as a starting point and price your own mix.

The budget tier, ranked by blended price

USD per 1M tokens · blended = (3 × input + 1 × output) ÷ 4 · list prices as of 2026-09-15 · verify with the provider.
ModelProviderInputOutputBlended
GPT-5 nano (legacy)OpenAI$0.05$0.40$0.14
Qwen-FlashAlibaba$0.05$0.40$0.14
Gemini 2.5 Flash-LiteGoogle$0.10$0.40$0.18
Mistral Small 4Mistral$0.15$0.60$0.26
DeepSeek V4.1 Flash (off-peak)DeepSeek$0.15$0.60$0.26
Llama 4 Scout (Together)Meta$0.18$0.59$0.28
GPT-5.6 LunaOpenAI$0.20$1.20$0.45
GPT-5.4 nanoOpenAI$0.20$1.25$0.46
Gemini 3.1 Flash-LiteGoogle$0.25$1.50$0.56
Gemini 3.5 Flash-LiteGoogle$0.30$2.50$0.85
GPT-5.4 miniOpenAI$0.75$4.50$1.69
Claude Haiku 4.5Anthropic$1.00$5.00$2.00

Blended prices are rounded to the cent. The tie at the top is new: Alibaba halved Qwen-Flash's input rate from $0.10 to $0.05 in early September 2026, taking it from $0.18 blended to $0.14 and level with legacy GPT-5 nano. That makes Qwen-Flash the current-generation floor outright, with Gemini 2.5 Flash-Lite — unchanged at $0.10 / $0.40 — now a step behind at $0.18. DeepSeek moved twice in the same window: it split its pricing into peak and off-peak bands on 2026-08-16, with peak hours billed at twice the off-peak rate, then cut the budget tier on 2026-09-10, when DeepSeek-V4.1-Flash replaced V4-Flash at $0.15 / $0.60 off-peak, down from $0.22 / $0.66. That takes it from $0.33 blended to $0.26, level with Mistral Small 4 — and off-peak it has the cheapest cached input in the tier, covered below.

The cheapest overall: a legacy model and a current one, tied

On raw price, legacy GPT-5 nano and Qwen-Flash now tie at $0.05 / $0.40 — a blended $0.14 each. Until September 2026 that floor belonged to GPT-5 nano alone, and the asterisk on it was "legacy": OpenAI has deprecated GPT-5 nano and scheduled its API shutdown for December 2026, so anything built on it must migrate within months. Qwen-Flash carries no such clock, which makes the tie lopsided in practice — the same sticker price, without the migration deadline. Among current-generation, generally-available budget models, the floor is now Qwen-Flash at $0.14, with Gemini 2.5 Flash-Lite next at $0.18.

What that means in dollars

Take a high-volume workload of 10M input and 2M output tokens per month. GPT-5 nano and Qwen-Flash both run about $1.30; Gemini 2.5 Flash-Lite about $1.80; DeepSeek V4.1 Flash off-peak about $2.70. The same workload on GPT-5.4 mini is about $16.50, and on Claude Haiku 4.5 about $20.00 — a 10-to-15× spread inside what everyone calls the "budget" tier. Picking the wrong budget model costs more than picking a flagship for a task that needed one.

DeepSeek still has the cheapest automatic cached input

If your workload reuses a large, stable prompt prefix, DeepSeek V4.1 Flash is the cheapest hands-off cache: cache hits list at $0.003 per 1M tokens off-peak ($0.006 peak) — a 98% discount off its $0.15 input rate, applied automatically with nothing to manage. Off-peak, that now undercuts the two models that used to read cheaper on paper. Legacy GPT-5 nano is $0.005, and it retires in December 2026. Qwen-Flash is also $0.005, but only on its explicit cache, which you pay $0.063 per 1M tokens to create and then have to manage; its automatic implicit cache reads at $0.01. The catch on DeepSeek is the same off-peak caveat: during peak hours the rate doubles to $0.006, just above both. For a cache-heavy pipeline scheduled off-peak, it is the one to beat.

The caveats that decide the real winner

  • Legacy status: GPT-5 nano is a prior-generation model — cheapest on sticker, but subject to deprecation. Weigh migration risk for anything long-lived.
  • Off-peak vs peak: DeepSeek's $0.15 / $0.60 (and its $0.003 cached rate) are off-peak; the peak rate (09:00–12:00 and 14:00–18:00 Beijing time, UTC+8) is 2×. These are the DeepSeek-V4.1-Flash rates in effect since 2026-09-10.
  • Third-party hosting: Meta does not sell Llama tokens directly — Llama 4 Scout's price is Together AI's serverless rate, and other hosts price it differently.
  • Tiered and promotional pricing: Qwen-Flash's $0.05 / $0.40 is the Singapore-region rate for prompts up to 256K input tokens; above that it bills $0.25 / $2.00, a 5× step. Qwen's budget rates can also include limited-time promos. Verify both the tier and the current price before you rely on them.
  • Audio surcharge: Gemini Flash-Lite models price audio input higher than text — Gemini 2.5 Flash-Lite text input is $0.10, but audio runs higher. Check the modality you actually send.
  • Cost per task, not per token: the cheapest per-token model is not automatically the cheapest per task — output length and tokenizer differences can flip the ranking. Model your real prompts.

Which should you pick

For most teams the honest answer is one of three:

  • Absolute lowest cost, short-lived or internal: legacy GPT-5 nano ($0.05 / $0.40), accepting deprecation risk — though Qwen-Flash now matches it without one.
  • Lowest cost on a current, GA model: Qwen-Flash ($0.05 / $0.40) after its September 2026 cut, with Gemini 2.5 Flash-Lite ($0.10 / $0.40) the next step up; DeepSeek V4.1 Flash off-peak wins for cache-heavy pipelines you would rather not hand-manage.
  • Lowest cost with no off-peak clock and a US-hosted API: GPT-5.6 Luna ($0.20 / $1.20) or Mistral Small 4 ($0.15 / $0.60).

What is the cheapest LLM API right now?

By blended price, legacy GPT-5 nano and Qwen-Flash tie at about $0.14 ($0.05 / $0.40) — but GPT-5 nano retires in December 2026, so among current-generation GA models Qwen-Flash is cheapest on its own, after Alibaba halved its input rate in early September 2026. Gemini 2.5 Flash-Lite follows at about $0.18 blended ($0.10 / $0.40); DeepSeek V4.1 Flash, after DeepSeek's 2026-09-10 cut, sits level with Mistral Small 4 at about $0.26 ($0.15 / $0.60 off-peak).

Is a cheaper model always the right choice?

No. A budget model that needs three attempts or a longer prompt to match a workhorse's single-shot answer can cost more in total, and it may miss on quality, latency, or tool use. Cheapest per token is a starting filter, not the decision.

How do I compute the cost for my own workload?

Take your monthly input and output token volumes, multiply each by the model's per-1M rate, and add them — then layer in cache-hit rate and any off-peak or long-context surcharge. The pricing calculator does this for any of the 40 models and lets you compare two side by side.

Do these prices change?

Frequently. These are list prices as of 2026-09-15; budget pricing is the most volatile tier — DeepSeek split its pricing into peak and off-peak bands on 2026-08-16, then cut its budget tier on 2026-09-10 — and promos and off-peak windows come and go. Re-verify with each provider before committing.

Cheap per token still adds up at volume. CloudQuell tracks your real OpenAI and Anthropic spend — per model, per token type — beside the rest of your cloud bill, at a flat monthly fee that never scales with usage.

Track your AI spend
← Back to all posts