This is an LLM API price comparison across the eight providers that matter for production traffic — OpenAI, Anthropic, Google, xAI, DeepSeek, Alibaba (Qwen), Mistral, and Meta (Llama) — organized by tier so you compare flagship against flagship and workhorse against workhorse. The sticker prices are one screen of numbers; the caveats underneath them are what actually decide the winner, so both are here.
All prices are USD per 1M tokens. "Cached input" is the discounted rate for reading a previously cached prompt prefix. Where a value is not published in the source dataset it is shown as "—".
Flagship tier
The frontier models, plus each provider's headline offering. Note the spread between Claude Fable 5.1 and GPT-6 Astra at the top and DeepSeek V4 Pro at the bottom — roughly 15× on input, 25× on output. These are not substitutes on quality, but the price gap is real.
| Model | Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|---|
| GPT-6 Astra (limited access at launch) | OpenAI | $10.00 | $50.00 | $1.00 | 1.05M |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | $0.25 | 1M |
| Claude Mythos 5.1 (trusted access, US only) | Anthropic | $10.00 | $50.00 | $0.25 | 1M |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $1.00 | 1M |
| GPT-5.6 Sol (cut rate to 2026-11-21) | OpenAI | $4.00 | $20.00 | $0.40 | 1.05M |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | $0.50 | 1M |
| Gemini 3.1 Pro (preview) | $2.00 | $12.00 | $0.20 | 1M | |
| Grok 4.6 | xAI | $2.00 | $6.00 | $0.50 | 500K |
| Grok 4.5 | xAI | $2.00 | $6.00 | $0.30 | 500K |
| Mistral Medium 3.5 | Mistral | $1.50 | $7.50 | — | — |
| Qwen-Max (≤32K) | Alibaba | $1.20 | $6.00 | — | 128K |
| DeepSeek V4 Pro (off-peak) | DeepSeek | $0.66 | $1.98 | $0.022 | 1M |
Read the flagship tier as four bands: the new top pair, where GPT-6 Astra and Claude Fable 5.1 both list $10.00 / $50.00 — identical to the cent, and separate only on cache reads, where Fable 5.1 and Mythos 5.1 price hits at 0.025× input ($0.25) against Astra’s $1.00; frontier Western models ($4 input and up — GPT-5.6 Sol on a cut rate, Claude Opus 5), value flagships ($1.20–$2.00 input — Gemini 3.1 Pro, Grok 4.6 and 4.5, Mistral Medium 3.5, Qwen-Max); and DeepSeek V4 Pro alone at the bottom on price. Fable 5.1 and Mythos 5.1 arrived on 2026-09-01 as one model at two safeguard levels — Mythos 5.1 is restricted to vetted cybersecurity and life-sciences organizations, currently US-only, so most teams are pricing Fable 5.1. Astra arrived on 2026-09-03 and is still rolling out — enterprises in OpenAI's Trusted Access Program first, with broader API access following — so treat its row as a price you can plan against rather than one you can necessarily buy today. The caveats section below explains why the cheapest rows come with strings attached.
Gemini vs Grok: the two value flagships
Gemini 3.1 Pro and Grok 4.5 both list $2.00 input, so head-to-head the difference is output: Grok 4.5 is $6.00 against Gemini 3.1 Pro's $12.00 — half the output cost — which makes Grok materially cheaper for generation-heavy work. Two caveats reorder that. Gemini 3.1 Pro is preview-priced and can change at general availability, and it adds a long-context surcharge above 200K tokens (input 2×, output 1.5×). Grok 4.5 doubles both input and output above 200K tokens and tops out at a 500K window versus Gemini's 1M. Under 200K tokens and output-heavy, Grok wins on price; if you need the 1M window, Gemini is the one that has it. xAI's newer Grok 4.6 keeps the same $2.00 / $6.00 list rates, 500K window, and 200K surcharge threshold (cached input rises to $0.50), so the comparison holds for it too.
Workhorse (mid) tier
This tier carries most production traffic — good enough for the majority of tasks, priced for volume. It is also the most competitive, with a 5× spread on input alone.
| Model | Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|---|
| GPT-5.4 | OpenAI | $2.50 | $15.00 | $0.25 | 1.05M |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $0.20 | 1M |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | $0.20 | 1.05M |
| Grok 4.3 | xAI | $1.25 | $2.50 | $0.20 | 1M |
| Gemini 3.6 Flash (promo to 2026-12-31) | $0.75 | $3.75 | $0.075 | 1M | |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | — | — |
| Qwen-Plus | Alibaba | $0.40 | $1.60 | — | 256K |
| Llama 4 Maverick (Together) | Meta | $0.27 | $0.85 | — | 1.05M |
Claude Sonnet 5's $2 / $10 began as an introductory price, but Anthropic made it the permanent standard rate in August 2026 — the scheduled increase to $3 / $15 was cancelled. Gemini 3.6 Flash is promo-priced at $0.75 / $3.75 through 2026-12-31 and lists $1.50 / $7.50 from 2027. Grok 4.3 is the standout on output at $2.50, a fraction of GPT-5.4's $15.00 — for chat and generation workloads where output dominates, it is the cheapest capable mid-tier model on a US-hosted API. Qwen-Plus bills output at $4.80 in thinking mode; the $1.60 shown is the standard rate.
Reasoning models: you pay for the thinking
Dedicated reasoning models spend extra output tokens "thinking" before they answer, so their output rate — and how many output tokens they burn — matters more than input. OpenAI sells these as a separate line; Anthropic and Google fold extended thinking into their main models, where the thinking tokens bill at the standard output rate.
| Model | Provider | Input | Output |
|---|---|---|---|
| o4-mini (retiring Oct 2026) | OpenAI | $1.10 | $4.40 |
| o3 (retiring Dec 2026) | OpenAI | $2.00 | $8.00 |
| o3-pro (retiring Dec 2026) | OpenAI | $20.00 | $80.00 |
| GPT-5.5 Pro | OpenAI | $30.00 | $180.00 |
o4-mini and o3 are the practical reasoning workhorses — though OpenAI has deprecated the whole o-series, with shutdowns in October and December 2026 and GPT-5.6 Sol's reasoning mode as the designated replacement. o3-pro and GPT-5.5 Pro are maximum-effort variants whose output rates ($80–$180) make them expensive to let run long. Because reasoning models emit far more output tokens than a normal chat turn, their effective cost per task can exceed a flagship's even when the per-token rate looks moderate.
The budget tier, in one line
Below the workhorse tier sits a field of sub-dollar models — GPT-5.6 Luna, GPT-5.4 nano, Gemini Flash-Lite, DeepSeek V4.1 Flash, Qwen-Flash, Mistral Small 4, Llama 4 Scout, and legacy GPT-5 nano — where the cheapest options run 10–15× less than the priciest "budget" model. That tier has its own ranked matrix; see the cheapest-LLM guide linked below.
The caveats that reorder the ranking
Every cheap row above carries a footnote. Read these before you rank on price:
- Preview pricing: Gemini 3.1 Pro is priced as a preview and can change at general availability.
- Off-peak vs peak: DeepSeek's rates are off-peak; the peak rate (09:00–12:00 and 14:00–18:00 Beijing time, UTC+8) is 2× on both V4 Pro and V4.1 Flash. DeepSeek introduced that peak/off-peak split on 2026-08-16, then cut its budget tier on 2026-09-10 when DeepSeek-V4.1-Flash replaced V4-Flash — older figures still circulate; the table above is current.
- Long-context surcharge: Gemini 3.1 Pro doubles input and adds 50% to output above 200K tokens; Grok 4.6 and 4.5 double input and output above 200K.
- Long-context surcharge, continued: GPT-6 Astra bills 2× input and cached input and 1.5× output once a prompt passes 272K tokens, applied to the entire request — $20.00 / $75.00 rather than $10.00 / $50.00.
- Tiered pricing: Qwen-Max bills by prompt length — $1.20 / $6.00 up to 32K input tokens, $2.40 / $12.00 from 32K to 128K (Singapore-region list prices; other regions differ).
- Promotions: OpenAI cut GPT-5.6 Sol from $5.00 / $30.00 to $4.00 / $20.00 (cache $0.40) on 2026-08-21 and guarantees the reduced rate at least through 2026-11-21; it has not published what happens after that date, so do not assume either the cut or the old price persists. Gemini 3.6 Flash bills $0.75 / $3.75 (cache $0.075) only through 2026-12-31, then $1.50 / $7.50 from 2027-01-01.
- Third-party hosting: Meta does not sell Llama tokens directly; the Llama prices shown are Together AI's serverless rates, and other hosts differ.
- Tokenizer skew: per-token prices are not comparable across tokenizer generations — Anthropic's newest models emit roughly 30% more tokens for the same text. Compare cost per task, not price per token.
Which should you pick
Match the model to the job, not to the top of a leaderboard:
- Cheapest capable flagship: DeepSeek V4 Pro off-peak ($0.66 / $1.98) still sits far below the Western flagships — if its peak-hours 2× and China-hosted API fit your constraints.
- Best value flagship on a US-hosted API: Grok 4.6 or 4.5 ($2.00 / $6.00) for output-heavy work; Gemini 3.1 Pro ($2.00 / $12.00) when you need the 1M window and can accept preview pricing.
- Frontier quality: GPT-5.6 Sol (promotional $4.00 / $20.00) undercuts Claude Opus 5 ($5.00 / $25.00); Claude Fable 5.1 ($10.00 / $50.00, cache reads $0.25) when you want Anthropic's most capable model and cost is secondary.
- Workhorse default: Claude Sonnet 5 ($2 / $10, now the standard price) or promo-priced Gemini 3.6 Flash; Grok 4.3 ($1.25 / $2.50) is the cheapest of the mid tier for output-heavy traffic.
What is the cheapest LLM API?
On sticker price, DeepSeek V4 Pro off-peak ($0.66 / $1.98) is the cheapest capable flagship, and in the budget tier legacy GPT-5 nano ($0.05 / $0.40) is the cheapest overall. But "cheapest" depends on your input:output ratio, whether you hit peak-hours surcharges, and how your text tokenizes — price your own workload rather than trusting a single headline number.
Are the prices the same across providers for the same quality?
No. At the flagship tier the input rate ranges from $0.66 (DeepSeek V4 Pro off-peak) to $10.00 (Claude Fable 5) — about a 15× spread — and the models are not interchangeable on quality, latency, tool support, or data residency. Price is one axis; the caveats above are the others.
Why do some models cost more above a token threshold?
Serving very long contexts is more expensive, so several providers add a long-context surcharge. Gemini 3.1 Pro doubles input and adds 50% to output above 200K tokens; Grok 4.6 and 4.5 double both above 200K. If your prompts are long, the headline rate understates your bill — model the surcharge tier explicitly.
How current are these prices?
They are list prices as of 2026-09-15 and they move often — recently OpenAI cut its GPT-5.6 Sol flagship to a promotional $4 / $20 (in effect at least through 2026-11-21), DeepSeek split its pricing into peak and off-peak bands and then cut its budget tier, Anthropic made Claude Sonnet 5's intro price permanent, and xAI shipped Grok 4.6. Re-verify against each provider before you commit a budget.
Once you have picked models, CloudQuell tracks what you actually spend on OpenAI and Anthropic — per model, per workspace, per token type — beside the rest of your cloud bill, at a flat monthly fee that never scales with usage.
Track your AI spend →