Llama API pricing: every model, every rate
Llama is open-weight, so the same model is hosted by many providers at wildly different rates. Below is what each variant bills at here.
llama-3-8b. Rates checked 2026-09-16.Full rate card
26 billable models in total, 11 of them text models. On those text models: 0 carry a rate verified against the vendor's own published pricing, the rest are gateway rates only. USD per 1M tokens, checked 2026-09-16.
| Model | Input | Output | Out/In | Official (in / out) |
|---|---|---|---|---|
Llama-3.1-405B | $3 | $6 | 2.0× | not verified |
Meta-Llama-3.1-405B-Instruct | $3 | $6 | 2.0× | not verified |
llama-3.1-405b-instruct | $3 | $6 | 2.0× | not verified |
llama-3.2-90b-vision-instruct | $3 | $9 | 3.0× | not verified |
llama-3-70b | $2 | $4 | 2.0× | not verified |
llama-3.1-70b | $2 | $4 | 2.0× | not verified |
llama-3.1-70b-instruct | $2 | $2 | 1.0× | not verified |
meta-llama/llama-3.1-70b-instruct | $2 | $2 | 1.0× | not verified |
meta-llama/llama-4-maverick | $1.25 | $5 | 4.0× | not verified |
meta-llama/llama-4-scout | $1.25 | $5 | 4.0× | not verified |
llama-2-13b | $1 | $1 | 1.0× | not verified |
llama-2-70b | $1 | $1 | 1.0× | not verified |
llama-2-7b | $1 | $1 | 1.0× | not verified |
llama-3-8b | $1 | $1 | 1.0× | not verified |
llama-3.1-8b | $1 | $1 | 1.0× | not verified |
llama-3.2-11b-vision-instruct | $1 | $1 | 1.0× | not verified |
llama-3-sonar-large-32k-chat | $0.75 | $0.75 | 1.0× | not verified |
llama-3-sonar-small-32k-chat | $0.75 | $0.75 | 1.0× | not verified |
llama-3.2-3b-instruct | $0.5 | $0.25 | 0.5× | not verified |
llama-3.3-70b-instruct | $0.36 | $0.36 | 1.0× | not verified |
llama-3.1-70b-instruct-turbo | $0.25 | $1 | 4.0× | not verified |
llama-3.2-1b-instruct | $0.25 | $0.0625 | 0.2× | not verified |
llama-3.2-90b-vision | $0.17 | $0.1938 | 1.1× | not verified |
llama-3.1-8b-instruct | $0.125 | $0.5 | 4.0× | not verified |
llama-4-maverick | $0.07 | $0.35 | 5.0× | not verified |
llama-3.3-70b | $0.05 | $0.148 | 3.0× | not verified |
Output is billed at a multiple of input on most models. The Out/In column is that multiple — it matters more than the input price once output dominates your bill.
Across 11 priced text models the blended rate spans $0.198 (llama-3.3-70b) to $9 (Meta-Llama-3.1-405B-Instruct) per 1M tokens — a 45× spread. Picking the wrong tier is usually the single most expensive mistake here.
What Meta costs at three usage levels
Priced on llama-3.3-70b at $0.05 in / $0.148 out; llama-3-8b at $1 in / $1 out per 1M tokens, with a 5:1 input-to-output ratio — roughly what an
interactive workload produces.
| Usage level | Tokens per month (in / out) | Lowest-cost tierllama-3.3-70b | Median-rate modelllama-3-8b |
|---|---|---|---|
| Light pilot or side project | 5M / 1M | $0.40 | $6.00 |
| Working one product in production | 50M / 10M | $3.98 | $60.00 |
| Heavy high-volume pipeline | 500M / 100M | $39.80 | $600.00 |
The first column is this vendor's lowest-cost tier; the second is the model closest to its median rate, which is nearer what a real deployment bills. Swap in your own token split in the calculator.
Same budget, other vendors
3 models from other vendors priced within 35% of
llama-3-8b on a blended rate — the cheapest, the dearest and one in between.
This is the horizontal check a single-vendor pricing page cannot give you.
| Model | Vendor | In / out per 1M | vs this |
|---|---|---|---|
qwen3-30b-a3b-think | Alibaba | $0.1 / $1.2 | −35.0% |
doubao-lite-32k | ByteDance | $1 / $1 | same |
deepseek-v4-pro-0813 | DeepSeek | $0.66 / $1.98 | +32.0% |
Blended comparison (input + output per 1M tokens). Price is one axis — these are the models worth testing side by side at this budget, not interchangeable substitutes.
Why the output rate decides the bill
Across the 6 priced Meta text models, the output rate runs from 2.0× to 5.0× the input rate, with a median of 4.0×. Output tokens are what a chat or agent turn produces — on a typical workload they are the smaller half of the token count but the larger half of the bill.
Worked out: at a 4.0× multiple, a job whose token count is 80% input still spends 50% of its cost on the output it generates. Comparing vendors on the input rate alone hides that.
What Meta is good at
Llama 405B is the usual open-weight choice when you need on-prem-equivalent capability without a frontier price tag.
Other vendors
Every rate card here is built the same way, from the same price snapshot (2026-09-16), so the numbers are comparable across vendors:
Anthropic · OpenAI · Google · DeepSeek · Alibaba · xAI · Zhipu · Moonshot · MiniMax · ByteDance
FAQ
How much does Meta charge per million tokens?
Across the 11 priced text models in this catalogue, the blended rate (input plus output per 1M tokens) runs from $0.198 on llama-3.3-70b to $9 on Meta-Llama-3.1-405B-Instruct, as of 2026-09-16. The spread is what matters: choosing the wrong tier inside one vendor usually costs more than switching vendor.
What is the cheapest Meta model for high-volume work?
llama-3.3-70b is the lowest blended rate here at $0.05 input and $0.148 output per 1M tokens. On a Heavy workload — 500M input and 100M output tokens a month — that bills about $39.80. At the other end, the same workload on Meta-Llama-3.1-405B-Instruct costs about $2,100.00.
Is Meta cheaper through a gateway than buying direct?
We do not claim a discount on Meta. Its list pricing is not published in a stable machine-readable form, so every figure on this page is a gateway rate with no verified official price to compare against. Check the number in your dashboard before committing to a budget.
How does output pricing change a Meta bill?
Output is billed at a multiple of input on every model priced here, so the input rate alone never predicts the bill. A workload that is mostly input tokens still lands most of its cost on the output column once that multiple is applied. Price your own in/out split in the calculator rather than extrapolating from the headline input rate.
Are the Llama models here the cheapest way to run open-weight work?
Among open-weight families in this catalogue they are the lowest-priced, with the 70B-class tiers in the sub-$1 blended band. The gap to the cheapest closed models is small, so the decision usually comes down to whether you need the weights themselves.
See all 827 models → or use the cost calculator with your own token split.