Qwen API pricing: every model, every rate
Qwen is the largest non-US model family by parameter spread, from 0.6B up to Max tiers. Below is every billable variant.
qwen3.8-max. Rates checked 2026-09-16.Full rate card
88 billable models in total, 47 of them text models. On those text models: 0 carry a rate verified against the vendor's own published pricing, the rest are gateway rates only. USD per 1M tokens, checked 2026-09-16.
| Model | Input | Output | Out/In | Official (in / out) |
|---|---|---|---|---|
qwen2.5-32b | $10 | $30 | 3.0× | not verified |
qwen2.5-vl-72b-instruct | $8 | $24 | 3.0× | not verified |
qwen2.5-vl-32b-instruct | $4 | $12 | 3.0× | not verified |
qwen3-coder | $3 | $12 | 4.0× | not verified |
qwen2-vl-72b-instruct | $2.25 | $2.25 | 1.0× | not verified |
qwen2.5-72b-instruct | $2 | $6 | 3.0× | not verified |
qwen2.5-math-72b-instruct | $2 | $6 | 3.0× | not verified |
qwen3-235b-a22b-think | $1.5 | $6 | 4.0× | not verified |
qwen3.7-max | $1.25 | $3.75 | 3.0× | not verified |
qwen3.5-omni-flash | $1.1 | $6.65 | 6.0× | not verified |
qwen-72b | $1 | $1 | 1.0× | not verified |
qwen2.5-32b-instruct | $1 | $3 | 3.0× | not verified |
qwen2.5-coder-14b-instruct | $1 | $3 | 3.0× | not verified |
qwen2.5-coder-32b-instruct | $1 | $3 | 3.0× | not verified |
qwen2.5-vl-7b-instruct | $1 | $2.5 | 2.5× | not verified |
qwen3-vl-235b-a22b | $1 | $4 | 4.0× | not verified |
qwen3.8-max | $1 | $3 | 3.0× | not verified |
qwen3.8-max-0902 | $1 | $3 | 3.0× | not verified |
qwen-mt-plus | $0.9 | $2.7 | 3.0× | not verified |
qwen-max | $0.8 | $3.2 | 4.0× | not verified |
qwen-omni-turbo | $0.8 | $2.2504 | 2.8× | not verified |
qwen-vl-plus | $0.75 | $2.25 | 3.0× | not verified |
qwen3-coder-480b-a35b-instruct | $0.75 | $3.75 | 5.0× | not verified |
qwen3.6-max-preview | $0.65 | $3.9 | 6.0× | not verified |
qwen2.5-vl-3b-instruct | $0.6 | $1.8 | 3.0× | not verified |
qwen3-max | $0.6 | $3 | 5.0× | not verified |
qwen3-max-preview | $0.6 | $3 | 5.0× | not verified |
qwen2.5-14b-instruct | $0.5 | $1.5 | 3.0× | not verified |
qwen2.5-14b-instruct-1m | $0.5 | $1.5 | 3.0× | not verified |
qwen2.5-coder-7b-instruct | $0.5 | $1 | 2.0× | not verified |
qwen2.5-math-7b-instruct | $0.5 | $1 | 2.0× | not verified |
qwen3-coder-plus | $0.5 | $2.5 | 5.0× | not verified |
qwen3-next-80b-a3b | $0.5 | $2 | 4.0× | not verified |
qwen-plus-busikao | $0.4 | $1 | 2.5× | not verified |
qwen-vl-max | $0.4 | $1.6 | 4.0× | not verified |
qwen-mt-turbo | $0.35 | $0.9765 | 2.8× | not verified |
qwen3-235b-a22b | $0.35 | $1.4 | 4.0× | not verified |
qwen3.5-397b-a17b | $0.3 | $1.8 | 6.0× | not verified |
qwen3.6-27b | $0.3 | $1.8 | 6.0× | not verified |
Qwen/Qwen3-Reranker-8B | $0.28 | — | not verified | |
qwen-plus-character | $0.25 | $0.7 | 2.8× | not verified |
qwen2-vl-7b-instruct | $0.25 | $0.25 | 1.0× | not verified |
qwen2.5-7b-instruct | $0.25 | $0.5 | 2.0× | not verified |
qwen2.5-7b-instruct-1m | $0.25 | $0.5 | 2.0× | not verified |
qwen3.6-plus | $0.25 | $1.5 | 6.0× | not verified |
qwen3.8-27b | $0.25 | $1.5 | 6.0× | not verified |
qwen3-coder-30b-a3b-instruct | $0.225 | $1.125 | 5.0× | not verified |
qwen-plus | $0.2 | $0.6 | 3.0× | not verified |
qwen3-vl-235b-a22b-instruct | $0.2 | $0.8 | 4.0× | not verified |
qwen3-vl-235b-a22b-thinking | $0.2 | $2 | 10.0× | not verified |
qwen3.5-122b-a10b | $0.2 | $1.6 | 8.0× | not verified |
qwen3.5-plus | $0.2 | $1.2 | 6.0× | not verified |
qwen3.7-plus | $0.2 | $0.8 | 4.0× | not verified |
qwen3.6-35b-a3b | $0.1875 | $1.125 | 6.0× | not verified |
qwen3-14b | $0.175 | $0.7 | 4.0× | not verified |
qwen2.5-3b-instruct | $0.15 | $0.45 | 3.0× | not verified |
qwen3-0.6b | $0.15 | $0.6 | 4.0× | not verified |
qwen3-1.7b | $0.15 | $0.6 | 4.0× | not verified |
qwen3-4b | $0.15 | $0.6 | 4.0× | not verified |
qwen3-coder-flash | $0.15 | $0.75 | 5.0× | not verified |
qwen3.5-27b | $0.15 | $1.2 | 8.0× | not verified |
Qwen/Qwen3-Reranker-4B | $0.14 | — | not verified | |
qwen3.5-35b-a3b | $0.125 | $1 | 8.0× | not verified |
qwen3-235b-a22b-instruct-2507 | $0.115 | $0.46 | 4.0× | not verified |
qwen3-235b-a22b-thinking-2507 | $0.115 | $1.15 | 10.0× | not verified |
qwen3-30b-a3b | $0.1 | $0.4 | 4.0× | not verified |
qwen3-30b-a3b-instruct-2507 | $0.1 | $0.4 | 4.0× | not verified |
qwen3-30b-a3b-think | $0.1 | $1.2 | 12.0× | not verified |
qwen3-30b-a3b-thinking-2507 | $0.1 | $1.2 | 12.0× | not verified |
qwen3-vl-30b-a3b-instruct | $0.1 | $0.4 | 4.0× | not verified |
qwen3-vl-30b-a3b-thinking | $0.1 | $1.2 | 12.0× | not verified |
qwen3-vl-plus | $0.1 | $0.8 | 8.0× | not verified |
qwen3-8b | $0.09 | $0.36 | 4.0× | not verified |
qwen3-vl-8b-instruct | $0.09 | $0.35 | 3.9× | not verified |
qwen3-vl-8b-thinking | $0.09 | $1.05 | 11.7× | not verified |
qwen3-32b | $0.08 | $0.32 | 4.0× | not verified |
qwen3-vl-32b-instruct | $0.08 | $0.32 | 4.0× | not verified |
qwen3-vl-32b-thinking | $0.08 | $0.32 | 4.0× | not verified |
qwen3-next-80b-a3b-instruct | $0.075 | $0.6 | 8.0× | not verified |
qwen3-next-80b-a3b-thinking | $0.075 | $0.6 | 8.0× | not verified |
qwen3.8-flash | $0.075 | $0.235 | 3.1× | not verified |
qwen3-rerank | $0.05 | $0 | not verified | |
qwen3.5-flash | $0.05 | $0.2 | 4.0× | not verified |
qwen-flash | $0.025 | $0.2 | 8.0× | not verified |
qwen-turbo | $0.025 | $0.1 | 4.0× | not verified |
qwen-turbo-1101 | $0.025 | $0.1 | 4.0× | not verified |
qwen3-vl-flash | $0.025 | $0.2 | 8.0× | not verified |
Qwen/Qwen3-Reranker-0.6B | $0.0105 | — | not verified |
Output is billed at a multiple of input on most models. The Out/In column is that multiple — it matters more than the input price once output dominates your bill.
Across 47 priced text models the blended rate spans $0.125 (qwen-turbo) to $15 (qwen3-coder) per 1M tokens — a 120× spread. Picking the wrong tier is usually the single most expensive mistake here.
What Alibaba costs at three usage levels
Priced on qwen3.8-max at $1 in / $3 out per 1M tokens, with a 5:1 input-to-output ratio — roughly what an
interactive workload produces.
| Usage level | Tokens per month (in / out) | Monthly cost qwen3.8-max |
|---|---|---|
| Light pilot or side project | 5M / 1M | $8.00 |
| Working one product in production | 50M / 10M | $80.00 |
| Heavy high-volume pipeline | 500M / 100M | $800.00 |
This is the model we cover in depth for this vendor. Swap in your own token split in the calculator.
Same budget, other vendors
3 models from other vendors priced within 35% of
qwen3.8-max on a blended rate — the cheapest, the dearest and one in between.
This is the horizontal check a single-vendor pricing page cannot give you.
| Model | Vendor | In / out per 1M | vs this |
|---|---|---|---|
gpt-5.4-mini | OpenAI | $0.375 / $2.25 | −34.4% |
doubao-lite-128k | ByteDance | $2 / $2 | same |
hunyuan-code | Tencent | $1.75 / $3.5 | +31.2% |
Blended comparison (input + output per 1M tokens). Price is one axis — these are the models worth testing side by side at this budget, not interchangeable substitutes.
Why the output rate decides the bill
Across the 46 priced Alibaba text models, the output rate runs from 2.5× to 12.0× the input rate, with a median of 4.0×. Output tokens are what a chat or agent turn produces — on a typical workload they are the smaller half of the token count but the larger half of the bill.
Worked out: at a 4.0× multiple, a job whose token count is 80% input still spends 50% of its cost on the output it generates. Comparing vendors on the input rate alone hides that.
What Alibaba is good at
Qwen coder models are strong enough for agentic work and priced well below comparable US models.
Detailed pages for Alibaba models
The models we cover in depth, each with worked cost at three usage levels and a comparison against similarly priced alternatives.
- Qwen3.8 Max — $1 / $3 per 1M tokens
Other vendors
Every rate card here is built the same way, from the same price snapshot (2026-09-16), so the numbers are comparable across vendors:
Anthropic · OpenAI · Google · DeepSeek · xAI · Zhipu · Moonshot · Meta · MiniMax · ByteDance
FAQ
How much does Alibaba charge per million tokens?
Across the 47 priced text models in this catalogue, the blended rate (input plus output per 1M tokens) runs from $0.125 on qwen-turbo to $15 on qwen3-coder, as of 2026-09-16. The spread is what matters: choosing the wrong tier inside one vendor usually costs more than switching vendor.
What is the cheapest Alibaba model for high-volume work?
qwen-turbo is the lowest blended rate here at $0.025 input and $0.1 output per 1M tokens. On a Heavy workload — 500M input and 100M output tokens a month — that bills about $22.50. At the other end, the same workload on qwen3-coder costs about $2,700.00.
Is Alibaba cheaper through a gateway than buying direct?
We do not claim a discount on Alibaba. Its list pricing is not published in a stable machine-readable form, so every figure on this page is a gateway rate with no verified official price to compare against. Check the number in your dashboard before committing to a budget.
How does output pricing change a Alibaba bill?
Output is billed at a multiple of input on every model priced here, so the input rate alone never predicts the bill. A workload that is mostly input tokens still lands most of its cost on the output column once that multiple is applied. Price your own in/out split in the calculator rather than extrapolating from the headline input rate.
Which Qwen tier is cheapest for bulk classification work?
The turbo and small tiers, which sit in the sub-$1 blended band. At that rate a Heavy workload still costs real money, so the practical move is to route only the ambiguous inputs to a larger Qwen tier and let the cheap tier handle the rest.
See all 827 models → or use the cost calculator with your own token split.