There is no cheapest model — only a cheapest model for your token mix
Every price table you have read ranks models by input price, because that is the column vendors print first. Real invoices do not look like that. Output tokens cost a median of 4.0× the input rate across 209 models, and on agent workloads output is most of what you buy. Re-rank the same 209 models on three realistic token mixes and the top ten barely overlap.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
Three mixes, three different winners
Below is the same 209-model index ranked three times: once for a chat workload (long context in, short answer out), once for coding agents (short prompt, very long output), once for batch generation (almost pure output). The blended figure is what 1M tokens actually costs you at that ratio.
1. Chat and RAG — 3 parts input to 1 part output
| Model | Vendor | Input | Output | Blended |
|---|---|---|---|---|
| Gemma2 9B It | $0.01 | $0.01 | $0.01 | |
| Gemma 3 1B It | $0.01 | $0.03 | $0.015 | |
| Gemma 3 4B It | $0.0175 | $0.0525 | $0.0263 | |
| Gemini 1.5 Flash 8B | $0.01875 | $0.075 | $0.0328 | |
| Qwen Turbo | Alibaba | $0.025 | $0.1 | $0.0438 |
| Qwen Turbo 1101 | Alibaba | $0.025 | $0.1 | $0.0438 |
| Gemma 3 12B It | $0.035 | $0.105 | $0.0525 | |
| Gemini 1.5 Flash | $0.0375 | $0.15 | $0.0656 | |
| Gemini 1.5 Flash 002 | $0.0375 | $0.15 | $0.0656 | |
| Gemini 2.0 Flash Lite | $0.0375 | $0.15 | $0.0656 |
Median across all 209 models at this mix: $0.9 per 1M tokens.
2. Coding agents and tool loops — 1 part input to 3 parts output
| Model | Vendor | Input | Output | Blended |
|---|---|---|---|---|
| Gemma2 9B It | $0.01 | $0.01 | $0.01 | |
| Gemma 3 1B It | $0.01 | $0.03 | $0.025 | |
| Gemma 3 4B It | $0.0175 | $0.0525 | $0.0437 | |
| Gemini 1.5 Flash 8B | $0.01875 | $0.075 | $0.0609 | |
| Qwen Turbo | Alibaba | $0.025 | $0.1 | $0.0813 |
| Qwen Turbo 1101 | Alibaba | $0.025 | $0.1 | $0.0813 |
| Gemma 3 12B It | $0.035 | $0.105 | $0.0875 | |
| DeepSeek OCR | DeepSeek | $0.108 | $0.108 | $0.108 |
| DeepSeek Ocr1 | DeepSeek | $0.108 | $0.108 | $0.108 |
| Gemini 1.5 Flash | $0.0375 | $0.15 | $0.1219 |
Median across all 209 models at this mix: $1.425 per 1M tokens.
3. Batch generation — 1 part input to 9 parts output
| Model | Vendor | Input | Output | Blended |
|---|---|---|---|---|
| Gemma2 9B It | $0.01 | $0.01 | $0.01 | |
| Gemma 3 1B It | $0.01 | $0.03 | $0.028 | |
| Gemma 3 4B It | $0.0175 | $0.0525 | $0.049 | |
| Gemini 1.5 Flash 8B | $0.01875 | $0.075 | $0.0694 | |
| Qwen Turbo | Alibaba | $0.025 | $0.1 | $0.0925 |
| Qwen Turbo 1101 | Alibaba | $0.025 | $0.1 | $0.0925 |
| Gemma 3 12B It | $0.035 | $0.105 | $0.098 | |
| DeepSeek OCR | DeepSeek | $0.108 | $0.108 | $0.108 |
| DeepSeek Ocr1 | DeepSeek | $0.108 | $0.108 | $0.108 |
| Mimo V2.5 | Other | $0.07 | $0.14 | $0.133 |
Median across all 209 models at this mix: $1.65 per 1M tokens.
Median blended rate moves from $0.9 on the chat mix to $1.425 on the agent mix and $1.65 on the batch mix — the same 209 models, only the ratio changed. And only 7 of the top ten survive the move from chat to batch.
Where “cheapest” flips outright
These models move the most between the two extremes. Every one of them is a model with a flat or near-flat output rate — cheap-looking on an input-ranked table, genuinely cheaper once output dominates.
| Model | Vendor | Input | Output | Chat rank | Batch rank | Move |
|---|---|---|---|---|---|---|
| Doubao Lite 32K | ByteDance | $1 | $1 | 114 | 72 | +42 up |
| Doubao Lite 4K | ByteDance | $1 | $1 | 115 | 73 | +42 up |
| Doubao Lite 128K | ByteDance | $2 | $2 | 155 | 113 | +42 up |
| Doubao Pro 32K | ByteDance | $2 | $2 | 156 | 114 | +42 up |
| Doubao Pro 4K | ByteDance | $2 | $2 | 157 | 115 | +42 up |
| Gemma 7B It | $1 | $1 | 116 | 74 | +42 up | |
| Qwen 72B | Alibaba | $1 | $1 | 112 | 71 | +41 up |
| Moonshot V1 8K | Moonshot | $1.2 | $1.2 | 129 | 91 | +38 up |
Largest move between the two mixes: 42 places out of 209.
The input-price trap, quantified
Compare the rank you get by sorting on input price alone against the rank you get by sorting on the real agent mix. The gap is the size of the mistake:
| Model | Vendor | Input | Output | By input | By agent mix | Error |
|---|---|---|---|---|---|---|
| Moonshot V1 8K | Moonshot | $1.2 | $1.2 | 163 | 97 | +66 |
| Gemma 7B It | $1 | $1 | 150 | 85 | +65 | |
| Grok 2 1212 | xAI | $1.5 | $1.5 | 173 | 109 | +64 |
| Doubao Lite 32K | ByteDance | $1 | $1 | 146 | 83 | +63 |
| Doubao Lite 4K | ByteDance | $1 | $1 | 147 | 84 | +63 |
| Gemma2 27B It | $0.63 | $0.63 | 124 | 61 | +63 | |
| Qwen 72B | Alibaba | $1 | $1 | 142 | 82 | +60 |
| Seed Oss 36B Instruct | ByteDance | $0.6 | $6 | 108 | 164 | -56 |
Largest gap: 66 places. At the top of the index that is the difference between %s and %s.
If you pick the vendor first: every vendor's floor
Most teams do not choose from a flat list of 209 models — they have a vendor already and want to know where the bottom of that range is. This is it, with the width of each range next to it.
| Vendor | Models | Entry model | Entry blended | Top blended | Spread |
|---|---|---|---|---|---|
| 26 | Gemma2 9B It | $0.01 | $3.906 | 391× | |
| Alibaba | 51 | Qwen Turbo | $0.0813 | $5.263 | 65× |
| DeepSeek | 12 | DeepSeek OCR | $0.108 | $6.5 | 60× |
| Other | 7 | Mimo V2.5 | $0.1225 | $2.625 | 21× |
| OpenAI | 40 | GPT 5 Nano | $0.1563 | $71.25 | 456× |
| Zhipu | 6 | GLM 5.3 Flash | $0.2062 | $8.875 | 43× |
| xAI | 26 | Grok 4 1 Fast Non Reasoning | $0.2125 | $10 | 47× |
| ByteDance | 13 | Doubao 1 5 Lite 32K | $0.2625 | $12 | 46× |
| Meta | 1 | Llama 4 Maverick | $0.28 | $0.28 | 1× |
| MiniMax | 6 | MiniMax M2.5 | $0.4875 | $6.825 | 14× |
| Moonshot | 9 | Kimi K2 | $0.975 | $21.875 | 22× |
| Mistral | 1 | Dolphin3.0 R1 Mistral 24B | $1.2751 | $1.275 | 1× |
| Anthropic | 11 | Claude Haiku 4.5 | $2 | $30 | 15× |
Measure your own mix before you read any of this
Ten lines of code is all it takes, and it replaces every assumption above with your own number:
# 先量出你自己的 output 占比,再去看榜单。
# 这一步花 10 分钟,能省掉一次错误的选型。
import anthropic, collections
client = anthropic.Anthropic()
stats = collections.Counter()
for _ in range(200): # 采样 200 次真实调用就够稳
r = client.messages.create(
model="claude-sonnet-5", max_tokens=1024,
messages=[{"role": "user", "content": next_prompt()}],
)
stats["in"] += r.usage.input_tokens
stats["out"] += r.usage.output_tokens
total = stats["in"] + stats["out"]
print("output share: %.1f%%" % (100 * stats["out"] / total))
# 把这两个数代回本页的 blended 公式,而不是照抄别人的榜单。
Then blend the two candidate rates at your ratio:
blended = (input_rate * in_share + output_rate * out_share)
That one line is the whole method. If you skip it, you are comparing someone else's workload to your invoice.
FAQ
Which model is the cheapest per million tokens?
It depends on your input-to-output ratio, which is why no single answer is honest. On a chat mix (3:1 input-heavy) the lowest blended rate in our 209-model index is Gemma2 9B It at $0.01 per 1M tokens; on an agent mix (1:3) it is still Gemma2 9B It, but the models behind it shuffle because output pricing carries most of the bill. Rank one list by input price alone and you can be 66 places out on an agent workload.
Why does ranking by input price give the wrong answer?
Because output tokens cost more — a median of 4.0x the input rate across 209 models, and up to 12x. A model with a cheap input rate and an expensive output rate looks like a bargain on a price table and expensive on an invoice. The 66-place error in the table above is the size of that mistake on real workloads.
How do I work out my own token mix?
Sample your real traffic rather than guessing. Log input_tokens and output_tokens from the API usage object over a few hundred calls and take the ratio — the snippet above does it in ten lines. Once you have the ratio, blend each candidate model's two rates at that ratio before comparing; do not compare input rates and hope.
Does the ranking change if I use cached prompts?
Yes, and in the same direction but harder. A cache hit bills around 0.1x base input, which shrinks the input side of your bill and makes the output rate matter even more. If a large share of your input is cached, weight the output rate more heavily than the raw mix suggests.
Are these rates the vendor's official prices?
No — they are gateway rates we bill, in USD per 1M tokens, captured 2026-09-20. 13 of the 209 models can be cross-checked against the vendor's own published list rate and are marked verified; the rest are listed at face value with no comparison implied. Rates move with upstream promotions, so re-check before you commit a pipeline.
Why only 209 models when the catalogue has 827?
827 is the raw catalogue; 209 is what survives cleaning — dated snapshots, duplicate billing variants, preview and test entries, and rows with no usable price are all removed. Those 209 are the entries you can actually be billed on. Publishing the raw 827 would put six near-identical variants of one model in the same table and make every ranking meaningless.
Should I just pick the cheapest entry model per vendor?
Only if it does the job. The entry rate tells you the floor of a vendor's range, which is useful for a first estimate, but the spread column shows how wide that range is — one vendor's top model costs hundreds of times its entry model. Treat the floor as a budget boundary, not a recommendation.