AI API prices compared — official vs discounted

Every price below is listed twice: what the vendor charges directly, and what the same model costs through the discounted route. Verified 2026-09-16.

ModelOfficial (in / out)Gateway (in / out)Saving
Claude Opus 5
Anthropic
$5 / $25 $2.5 / $12.5 -50%
Claude Sonnet 5
Anthropic
$2 / $10 $1 / $5 -50%
Claude Fable 5
Anthropic
$10 / $50 $5 / $25 -50%
Claude Haiku 4.5
Anthropic
$1 / $5 $0.5 / $2.5 -50%
GPT-5.6 Sol
OpenAI
$4 / $20 $2.5 / $15 -27%
GPT-5.6 Terra
OpenAI
$2 / $12 $1 / $6 -50%
GPT-5.6 Luna
OpenAI
$0.2 / $1.2 $0.1 / $0.6 -50%
GPT-6 Astra
OpenAI
$10 / $50 $5 / $25 -50%
Gemini 3.7 Flash
Google
not verified $0.375 / $1.875 n/a
DeepSeek V4 Pro
DeepSeek
not verified $0.66 / $1.98 n/a
DeepSeek V4 Flash
DeepSeek
not verified $0.22 / $0.66 n/a
GLM-5.3
Zhipu
not verified $0.7 / $2.2 n/a
MiniMax M3
MiniMax
not verified $0.15 / $0.6 n/a
Kimi K3
Moonshot
not verified $1.5 / $7.5 n/a
Qwen3.8 Max
Alibaba
not verified $1 / $3 n/a
Grok 4.1
xAI
not verified $1 / $5 n/a

Per 1M tokens, USD, standard tier. Synced 2026-09-16. Savings are shown only where we could verify the official list price against vendor documentation — see each model page for the source. Where a cell reads not verified we make no comparison claim.

Check live rates on the platform →Free to sign up · $1 minimum top-up · No prepayment
Read this before you compare: the discount applies to US frontier models. DeepSeek, GLM, Kimi, Qwen and MiniMax are billed at their official list price — the point of listing them is access, not saving.

How to read the table

Pick a model

Claude Opus 5

Anthropic's strongest model for long-horizon agentic coding and enterprise work.

Claude Sonnet 5

The daily driver — near-flagship coding quality at a fifth of Opus's output rate.

Claude Fable 5

Frontier Claude for demanding reasoning and long agent runs.

Claude Haiku 4.5

Ultra-fast classification, extraction and tight loops.

GPT-5.6 Sol

OpenAI's heaviest 5.6 tier — deepest reasoning, widest tool use.

GPT-5.6 Terra

The middle tier most teams settle on.

GPT-5.6 Luna

Lowest-cost tier in the GPT-5.6 family — drafts, routing, high-frequency calls.

GPT-6 Astra

OpenAI's newest frontier tier.

Gemini 3.7 Flash

Google's fast multimodal model with a 1M context window.

DeepSeek V4 Pro

DeepSeek's flagship — strong reasoning at open-weight economics.

DeepSeek V4 Flash

Peak/off-peak pricing — schedule batch work off-peak to cut cost sharply.

GLM-5.3

Zhipu's current flagship.

MiniMax M3

MiniMax's latest — very low price for the capability.

Kimi K3

Moonshot's latest flagship — long-context agentic work.

Qwen3.8 Max

Alibaba's largest Qwen tier.

Grok 4.1

xAI's flagship model.

Head-to-head comparisons

Deciding between two specific models? These pages price both against each other — including the cases where the ranking changes once you use gateway rates instead of vendor list prices.

Claude vs GPT

Sonnet 5 against GPT-5.6 Terra on blended cost and monthly spend.

Opus 5 vs GPT-5.6 Sol

The pair where the list-price ranking flips at gateway rates.

Claude vs Gemini

The cheap tier, where volume amplifies every fraction of a cent.

GPT vs Gemini

The sub-$1 tier compared on real rates and monthly totals.

Opus 5 vs Sonnet 5

What the 2.5× multiplier costs, and the routing pattern that avoids it.

DeepSeek vs Claude

The widest price gap people actually compare, costed honestly.

Lowest-cost LLM APIs

Lowest blended rates in the index, plus each vendor's floor price.

Ways to pay for Claude

Direct, prepaid bundle or pay-as-you-go — with the prepay maths.

vs OpenRouter

Breadth versus per-token cost, and when to run both.

Get access at these prices