The 2026 AI API price index: 209 models, priced the way you actually get billed
Every public price table ranks models by input price, because that is the one column every vendor prints. It is also the column that least resembles your invoice. This index re-ranks the same 209 models by three realistic token mixes, publishes the arithmetic, and gives you the whole dataset as CSV and JSON.
What this is
A benchmark of 209 billable models across 13 vendors, measured on 2026-09-20 from a single live rate configuration of 828 raw catalogue entries. Every number on this page is computed from that dataset at build time — nothing here is typed in by hand, and nothing is estimated.
The three mixes we use: 3:1 (chat — short questions, long answers), 1:3 (agents and coding — large context in, longer generation out), and 1:9 (batch — small prompt, bulk output). A blended rate is just (input×a + output×b) / (a+b).
Finding 1 — The ranking depends on your workload, not on the model
Sort by input price and sort by a realistic 1:3 agent mix and you get two different lists. The disagreement is not marginal: across 209 models the same model moves by up to 66 places. Re-rank for a 1:9 batch mix and the largest single move is 42 places.
Of the three top-ten lists — chat, agent, batch — only 7 models appear in all three. The union is 13 models. There is no winner across workloads, only a winner per workload.
| Model | Vendor | In /1M | Out /1M | Chat 3:1 | Agent 1:3 | Batch 1:9 | Shift |
|---|---|---|---|---|---|---|---|
| Doubao Lite 32K | ByteDance | $1 | $1 | #114 | #83 | #72 | +42 |
| Doubao Lite 4K | ByteDance | $1 | $1 | #115 | #84 | #73 | +42 |
| Doubao Lite 128K | ByteDance | $2 | $2 | #155 | #125 | #113 | +42 |
| Doubao Pro 32K | ByteDance | $2 | $2 | #156 | #126 | #114 | +42 |
| Doubao Pro 4K | ByteDance | $2 | $2 | #157 | #127 | #115 | +42 |
| Gemma 7B It | $1 | $1 | #116 | #85 | #74 | +42 | |
| Qwen 72B | Alibaba | $1 | $1 | #112 | #82 | #71 | +41 |
| Moonshot V1 8K | Moonshot | $1.2 | $1.2 | #129 | #97 | #91 | +38 |
| Grok 2 1212 | xAI | $1.5 | $1.5 | #138 | #109 | #104 | +34 |
| Moonshot V1 32K | Moonshot | $2.4 | $2.4 | #164 | #135 | #132 | +32 |
Across 209 billable AI models, ranking by input price instead of by a realistic 1:3 input-to-output mix moves a model by up to 66 places out of 209 — aitokensai, AI API Price Index 2026 (2026-09-20).
The practical consequence: a "lowest-cost model" claim without a stated token mix is not a comparison, it is a coincidence of column choice.
Finding 2 — Output tokens are the bill
Median input price across the set is $0.55 per 1M. Median output price is $1.8. Per model, the median output-to-input ratio is 4.0× (mean 4.5×, steepest 12×), and 16 models bill input and output at an identical rate.
At the median ratio on a 1:3 mix, output is 75% of the tokens but about 92% of the cost. Trimming a prompt feels productive and moves the invoice very little; capping generation length, asking for bounded structured output, or caching the stable prefix does the actual work.
| Model | Vendor | In /1M | Out /1M | Ratio |
|---|---|---|---|---|
| Gemma2 9B It | $0.01 | $0.01 | ×1.0 | |
| Gemma 3 1B It | $0.01 | $0.03 | ×3.0 | |
| Gemma 3 4B It | $0.0175 | $0.0525 | ×3.0 | |
| Gemini 1.5 Flash 8B | $0.01875 | $0.075 | ×4.0 | |
| Qwen Turbo | Alibaba | $0.025 | $0.1 | ×4.0 |
| GPT 5.4 Pro | OpenAI | $15 | $90 | ×6.0 |
| GPT 5.5 Pro | OpenAI | $15 | $90 | ×6.0 |
| GPT 5.2 Pro | OpenAI | $10.5 | $84 | ×8.0 |
Output tokens are the bill. Median output-to-input price ratio across 209 models is 4.0x (mean 4.5x, max 12x); at a 1:3 mix output carries about 92% of the cost while being 75% of the tokens — aitokensai, AI API Price Index 2026 (2026-09-20).
Finding 3 — The spread inside one vendor dwarfs the spread between vendors
OpenAI carries 40 billable models whose 1:3 blended rate runs from $0.15625 (GPT 5 Nano) to $71.25 (GPT 5.4 Pro) — a 456× spread inside a single price list. Google spans 391×. Choosing the wrong tier inside a vendor you already use is routinely more expensive than switching vendors.
| Vendor | Models | Lowest 1:3 | Median 1:3 | Highest 1:3 | Spread |
|---|---|---|---|---|---|
| 26 | $0.01 | $0.41875 | $3.90625 | ×391 | |
| Alibaba | 51 | $0.08125 | $0.85 | $5.2625 | ×65 |
| DeepSeek | 12 | $0.108 | $1.09609 | $6.5 | ×60 |
| Other | 7 | $0.1225 | $0.875 | $2.625 | ×21 |
| OpenAI | 40 | $0.15625 | $4.32812 | $71.25 | ×456 |
| Zhipu | 6 | $0.20625 | $1.825 | $8.875 | ×43 |
| xAI | 26 | $0.2125 | $1.3375 | $10 | ×47 |
| ByteDance | 13 | $0.2625 | $2 | $12 | ×46 |
| Meta | 1 | $0.28 | $0.28 | $0.28 | ×1 |
| MiniMax | 6 | $0.4875 | $1.95 | $6.825 | ×14 |
| Moonshot | 9 | $0.975 | $2.4 | $21.875 | ×22 |
| Mistral | 1 | $1.27505 | $1.27505 | $1.27505 | ×1 |
| Anthropic | 11 | $2 | $10 | $30 | ×15 |
Within a single vendor, per-token cost spans up to 456x (OpenAI: $0.15625 to $71.25 on a 1:3 mix). Picking the wrong model costs more than picking the wrong vendor — aitokensai, AI API Price Index 2026 (2026-09-20).
Finding 4 — Entry thresholds differ by 200×
The cheapest way to start differs enormously by vendor. The lowest 1:3 blended rate in the set is $0.01 (Gemma2 9B It, Google). The highest entry point is $2 (Claude Haiku 4.5, Anthropic) — 200× more expensive before you have written a line of code. That gap is a statement about product positioning, not about quality: the expensive entry buys a capability floor, not a price premium.
The cheapest entry point in this dataset costs $0.01 per 1M tokens on a 1:3 mix; the most expensive costs $2 — a 200x difference before a single request is sent.
Finding 5 — The distribution is bottom-heavy
81 of 209 models (39%) land under $1 per 1M tokens on a 1:3 mix. The median is $1.425. Only 16 models (8%) exceed $10. The expensive tail is real but thin, and it is where nearly all of the public conversation happens.
| 1:3 blended rate | Models | Share |
|---|---|---|
| under $0.5 | 56 | 27% |
| $0.5 – $1 | 25 | 12% |
| $1 – $3 | 61 | 29% |
| $3 – $10 | 45 | 22% |
| $10 – $50 | 19 | 9% |
| over $50 | 3 | 1% |
Finding 6 — Most of this cannot be checked against a published rate
Of 209 models, only 13 have a published list price we can hold against the gateway rate. Those rows are marked verified. The rest are labelled as gateway rates and are not presented as discounts, because there is nothing to discount from. Any site that shows you a percentage saving across hundreds of models is comparing against a number it invented.
Method
- Source. One gateway's live rate configuration, 828 raw catalogue entries, read at build time.
- Cleaning. 209 entries survive. Removed: date-stamped snapshots, duplicate
effort/reasoning variants that resolve to one model, test and preview builds, entries with no
price, and non-conversation endpoints. The rules live in
tools/clean_catalog.pyand are the same rules used by every other page on this site — the calculator, the catalogue and this index cannot drift apart. - Normalisation. All prices in USD per 1M tokens.
- Derivation. Blended rates are arithmetic on the two published prices. No cached-token, batch or volume assumption is folded in.
- Known gaps. No caching, no batch discount, no context-window cost, no image/audio token pricing, no per-request fees.
Download the data
The full 209-row dataset is published under CC BY 4.0. Reuse is fine with attribution and a link; please link to the canonical URL rather than mirroring the numbers, because a stale copy helps nobody.
Columns: model id, vendor, display name, input and output price per 1M, three blended rates, official list price where one exists, and a verified flag.
How to cite
If you use the numbers, here is the attribution already written out:
aitokensai. AI API Price Index 2026: 209 models benchmarked by real token mix. aitokensai. 2026-09-20. https://aitokensai.com/benchmark/price-index-2026/ (CC BY 4.0)
Canonical page: https://aitokensai.com/benchmark/price-index-2026/
· Data: https://aitokensai.com/benchmark/price-index.csv ·
Licence: CC BY 4.0
Changelog
- 2026-09-20 — first publication. 209 models, 13 vendors, 13 verified against a published list rate.
FAQ
Is this every model that exists?
No. It is every billable entry in one gateway's live rate configuration: 828 raw entries reduced to 209 after removing date-stamped snapshots, duplicate effort variants, test and preview builds, and entries with no price. Treat it as a large, consistently measured sample rather than a census of the industry.
Are these the prices vendors publish?
For 13 models, yes — we hold a published list rate and mark those rows verified. For the rest we show the gateway rate and say so on the page. We never publish a derived number as if it were an official one, and we make no claim of a flat discount.
What is a blended rate?
A weighted average of input and output price for a stated token mix, (input×a + output×b) / (a+b). We use 3:1 for chat, 1:3 for agent and coding workloads, and 1:9 for batch. It is arithmetic on published prices, not a new charge.
What is not included?
Prompt caching, batch discounts, context-window costs, image and audio tokens, and any per-request fee. Caching in particular can cut input cost by an order of magnitude on stable prompts — see the prompt caching guide for the break-even arithmetic.
Can I use the data?
Yes. The CSV and JSON below are published under CC BY 4.0 — reuse is fine with attribution and a link. Re-hosting a stale copy is not useful to anyone, so link to the canonical URL rather than mirroring the numbers.
Related
- Lowest-cost model by workload — the full three-way ranking behind Finding 1
- Output token pricing — the ratio distribution behind Finding 2
- Why the same model has two prices — what "verified" does and does not mean here
- Prompt caching savings — the largest lever this index deliberately excludes
- Full catalogue — all 209 rows, filterable
- Cost calculator — model your own token mix