Skip to content
INFERENCE, SERVING AND COST CONTROL

LLM API Pricing 2026: Every Major Model Compared

I rebuild my cost spreadsheet monthly, because these numbers refuse to sit still. Here is where every major API stands as of August 3, 2026, checked against provider pricing pages that day. All figures are per million tokens, standard tier.

The Table, as of August 3, 2026

Model Input $/Mtok Output $/Mtok Context Best for
Claude Fable 5 $10.00 $50.00 1M Hardest reasoning, long agent runs
Claude Opus 5 $5.00 $25.00 1M Complex agentic coding
Claude Sonnet 5 $2.00* $10.00* 1M Production workhorse
Claude Haiku 4.5 $1.00 $5.00 200K Classification and routing
GPT-5.6 Sol $5.00 $30.00 1.05M OpenAI’s frontier tier
GPT-5.6 Terra $2.00 $12.00 1.05M Everyday coding and agents
GPT-5.6 Luna $0.20 $1.20 1.05M Very high volume, low latency
Gemini 3.1 Pro Preview $2.00 $12.00 1M Multimodal (over 200K: $4/$18)
Gemini 3.6 Flash $1.50 $7.50 1M Long-context agent loops
Gemini 3.1 Flash-Lite $0.25 $1.50 1M Bulk text processing
DeepSeek V4-Pro $0.435 $0.87 1M Near-frontier at open-weight prices
DeepSeek V4-Flash $0.14 $0.28 1M Cheapest useful model
Qwen3.7 Max $1.25** $3.75** 1M Strong non-US flagship
Qwen3.5 Flash $0.10 $0.40 check provider page Bulk classification

* Sonnet 5 introductory rate runs through August 31, 2026, then moves to $3.00/$15.00. ** Qwen3.7 Max is on a 50% promotion; list price is $2.50/$7.50.

What Jumps Out

Output tokens dominate. Every provider charges 2x to 6x more for output than input, so a chatty model costs more than its input rate suggests. Fable 5 sits at 5x, GPT-5.6 Sol at 6x, DeepSeek at a flat 2x.

The spread top to bottom is roughly 70x on input, from Fable 5 at $10 to DeepSeek V4-Flash at $0.14. Prices also move mid-quarter: OpenAI cut Luna by about 80% and Terra by about 20% on July 30, three weeks after the GPT-5.6 family shipped. Anything you read from June is stale.

A Real Workload: 10M Tokens a Month

A support-triage service handling 8M input and 2M output tokens monthly, priced three ways.

Claude Opus 5: (8 × $5) + (2 × $25) = $90/month

GPT-5.6 Terra: (8 × $2) + (2 × $12) = $40/month

DeepSeek V4-Flash: (8 × $0.14) + (2 × $0.28) = $1.68/month

That last number is not a typo. DeepSeek has undercut US labs by 10x to 30x for two years, and V4 went GA on July 20 with a 1.6T-parameter Pro variant and native 1M context. For triage and extraction, the quality gap is small enough that the 50x price gap wins. My proprietary versus open-source breakdown covers where that stops working.

Pattern 1: Prompt Caching

Caching is the highest-leverage change most teams have not made. Anthropic charges 10% of the input rate on a cache hit. DeepSeek charges $0.0028 per million on a V4-Flash hit against $0.14 on a miss, 50x cheaper. Google and Qwen both advertise up to 90% off cached input.

The catch is that caching is a prefix match. Put your system prompt, tool definitions, and retrieved documents first, keep them byte-identical, volatile parts last. One timestamp near the top invalidates everything after it, and you pay full price wondering why the bill never dropped.

Pattern 2: Model Routing and Cascades

Here is what I route to Haiku and why. Anything with a fixed output shape goes there: intent classification, language detection, PII redaction, tool selection. Those tasks produce 30 output tokens, and a frontier model produces 30 identical tokens for 25x the price.

Cascades take this further. Run the cheap model first, escalate only when it flags low confidence or a validator rejects the output. A cascade resolving 80% of traffic on Luna and escalating the rest to Sol lands near $1.20 per million blended, not $5. I cover the confidence-threshold mechanics in my post on cutting inference costs by 80%.

Pattern 3: Batch APIs

Every major provider discounts asynchronous work by 50%: Anthropic, OpenAI, Google, Alibaba. Submit a job, results arrive within 24 hours, pay half.

Nightly embeddings, backfills, evaluation runs, summarization pipelines. None of that needs a response in 400 milliseconds. Combine batch with caching on a shared prefix and the effective rate drops below 20% of sticker. Teams skip this because it needs a queue, and the queue is forty lines of code.

What GPT-4-Level Intelligence Costs Now

The number I keep coming back to: GPT-4-level capability ran about $30 per million tokens in early 2023. Today it costs under $1, per llm-stats.com, which tracks roughly a 10x annual decline for equivalent capability.

That reframes the whole table. Fable 5 at $10 looks expensive against Haiku at $1, and in absolute terms it is. Against 2023 pricing for far weaker output, it is a bargain. The frontier tier stays roughly flat while the capability behind it climbs; everything below collapses toward zero within about eighteen months.

Practical consequence: do not architect around today’s prices. Build a routing layer, keep model IDs in config, re-benchmark quarterly. The model you cannot afford in August is often the default by March. My comparisons of GPT-5.3 Codex against Claude Opus 4.6 and Fable 5 against GPT-5.6 show how fast that ordering shifts.

Frequently Asked Questions

Which LLM API is cheapest in 2026?

DeepSeek V4-Flash at $0.14 input and $0.28 output, verified on DeepSeek’s pricing page on August 3, 2026. Qwen3.5 Flash is close at $0.10/$0.40, and its Beijing endpoint is cheaper still. Among US providers, GPT-5.6 Luna at $0.20/$1.20 is the floor. Add caching and batch and any of these drops another 50% to 90%.

Do I always save money by picking a cheaper model?

No. A weaker model that needs three retries or a bigger few-shot prompt can cost more than one strong call. Measure cost per resolved task, never cost per token. I have watched Haiku pipelines lose to Sonnet on total spend because the retry loop ate the savings.

Why is output so much more expensive than input?

Input tokens are processed in parallel during prefill, which is compute-efficient. Output tokens come out one at a time, each needing a full forward pass while the GPU sits mostly idle on memory bandwidth. That asymmetry is physical, and it shows up in every price sheet.

Does the 1M context window cost extra?

Sometimes. Gemini 3.1 Pro doubles input pricing above a 200K-token prompt, and GPT-5.6 raises all three tiers for long-context requests. Anthropic and DeepSeek serve their full 1M windows at standard rates. Check that before stuffing a whole repository into one call, and see my practical model comparison for how context size plays out.

How reliable is this table next month?

Treat it as a snapshot dated August 3, 2026. Terra and Luna changed price on July 30, Opus 5 launched July 24, and Sonnet 5’s introductory rate expires August 31. Confirm against the official Anthropic pricing page or your provider’s equivalent before committing to a budget.