LiteLLM model names: the prefix is a routing instruction
LiteLLM reads your model string as two things, not one. It splits at the first slash: the segment before names the provider, and everything after is the model id handed to that provider untouched. Get that split wrong and you get one of two failures that look alike but are opposites — one raised before any request is sent, one returned by a server that received a perfectly routed call.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
What the split actually does
LiteLLM's own documentation for OpenAI-compatible endpoints states the rule directly: put openai/ in front of your model name so LiteLLM knows you are calling a /chat/completions endpoint. For the older /completions route the documented prefix is text-completion-openai/ instead. The part after the first slash is passed to the provider unchanged.
"Unchanged" is the half people miss. It means additional slashes are not separators — they are part of the model id. So openai/glm/glm-5.2 resolves to provider openai and sends glm/glm-5.2 upstream. That is not a hack; it is the only way to reach an endpoint that registers its ids with a vendor segment already in them.
def split_model(model: str):
"""Mirror LiteLLM's documented rule: the first segment names the provider,
everything after the first slash is the model id and is passed through
to that provider unchanged -- including any further slashes."""
provider, sep, rest = model.partition("/")
if not sep:
return None, model # no provider in the string: routing must be inferred
return provider, rest
for m in ["gpt-5.6-luna",
"openai/gpt-5.6-luna",
"openai/glm/glm-5.2",
"anthropic/claude-sonnet-5",
"support-bot"]:
provider, sent = split_model(m)
if provider is None:
print(f"{m:30} -> no provider given, LiteLLM must guess or refuse")
else:
print(f"{m:30} -> provider={provider:10} sent upstream={sent}")
Run that against the ids you actually use before you debug anything else. It needs no LiteLLM installed, and it tells you in one line what the endpoint is going to receive.
The error that fires before the request does
Pass a string with no provider and no match in LiteLLM's internal map, and you get: litellm.exceptions.BadRequestError: LLM Provider NOT provided. Pass in the LLM provider you are trying to call. You passed model=support-bot.
Read the last clause literally — it is telling you the request never happened. No socket opened, no token counted, nothing billed. This is the single most useful fact on the page, because it splits the failure space cleanly:
- The routing error is free. It is raised inside your process. Retrying it, or rotating a key over it, cannot work — the request is identical and so is the answer.
- A 404 is not free and is not the same thing. It means routing succeeded, a request went out, and the endpoint rejected the id. Now the key is proven valid and the id is the suspect.
- A 401 is a third thing again. Authentication precedes routing on the server, so the id may be fine and the credential is not. Our error-code walkthrough covers that ordering.
Three ways to set the provider, and where each belongs
The prefix in the string is the most explicit and the easiest to read in a log. It is also the only one that survives being handed to a framework that re-serialises config.
import os
import litellm
# The prefix is a routing instruction, not part of the model name.
# LiteLLM strips the first segment and sends the rest to the endpoint as-is.
resp = litellm.completion(
model="openai/gpt-5.6-luna",
api_base="https://aicomp.ai/v1",
api_key=os.environ["GATEWAY_API_KEY"],
messages=[{"role": "user", "content": "Reply with the single word: ok"}],
max_tokens=8,
)
print("answered by:", resp.model) # identity, as the endpoint reports it
print("usage :", resp.usage) # None here means usage was never returned
# When you cannot change the model string (it comes from a database, a UI,
# a framework config), pass the provider as its own argument instead:
same = litellm.completion(
model="gpt-5.6-luna",
custom_llm_provider="openai", # same routing, no prefix in the string
api_base="https://aicomp.ai/v1",
api_key=os.environ["GATEWAY_API_KEY"],
messages=[{"role": "user", "content": "Reply with the single word: ok"}],
max_tokens=8,
)
But the right place for the mapping, once you have more than one caller, is the proxy. There, model_name is what callers ask for and litellm_params.model is what is sent — so a friendly alias is decoupled from the qualified id behind it.
# litellm_config.yaml
# Callers ask for a friendly name; the proxy owns the provider-qualified one.
# Swapping the backend later is a one-line change here, not a client change.
model_list:
- model_name: fast
litellm_params:
model: openai/deepseek-v4-flash-0731
api_base: https://aicomp.ai/v1
api_key: os.environ/GATEWAY_API_KEY
- model_name: strong
litellm_params:
model: openai/claude-sonnet-5
api_base: https://aicomp.ai/v1
api_key: os.environ/GATEWAY_API_KEY
That indirection is what makes a backend swap a one-line change instead of a coordinated release across every service that hardcodes a model name. It is also the natural place to put a fallback, and therefore the place to think about how many upstream calls one request can turn into — which is why the retry and timeout budget is the next page to read once names resolve.
How many ids actually need the stacked form
In the 839-model catalogue we track, 12 ids contain a slash. For those, a single openai/ prefix is not enough — the vendor segment has to be preserved, which is exactly what stacking does.
| Registered id | With the provider prefix | Sent upstream |
|---|---|---|
BAAI/bge-reranker-v2-m3 | openai/BAAI/bge-reranker-v2-m3 | BAAI/bge-reranker-v2-m3 |
Pro/BAAI/bge-reranker-v2-m3 | openai/Pro/BAAI/bge-reranker-v2-m3 | Pro/BAAI/bge-reranker-v2-m3 |
Qwen/Qwen3-Reranker-0.6B | openai/Qwen/Qwen3-Reranker-0.6B | Qwen/Qwen3-Reranker-0.6B |
Qwen/Qwen3-Reranker-4B | openai/Qwen/Qwen3-Reranker-4B | Qwen/Qwen3-Reranker-4B |
Qwen/Qwen3-Reranker-8B | openai/Qwen/Qwen3-Reranker-8B | Qwen/Qwen3-Reranker-8B |
anthropic/claude-3.7-sonnet | openai/anthropic/claude-3.7-sonnet | anthropic/claude-3.7-sonnet |
anthropic/claude-sonnet-4 | openai/anthropic/claude-sonnet-4 | anthropic/claude-sonnet-4 |
meta-llama/llama-3.1-70b-instruct | openai/meta-llama/llama-3.1-70b-instruct | meta-llama/llama-3.1-70b-instruct |
Showing 8 of 12. The remaining 827 ids carry no slash, so a single prefix is sufficient for them.
The general rule, and the only one worth memorising: never derive the id from anything except the endpoint's own list. Not from a docs page, not from a URL slug, not from what worked on a different endpoint.
# The endpoint's own list is the only authoritative source of model ids.
curl -s https://aicomp.ai/v1/models -H "Authorization: Bearer $GATEWAY_API_KEY" > /tmp/models.json
head -c 200 /tmp/models.json
# -> starts with {"object":"list" ... good
# -> starts with <!DOCTYPE or <html something else answered; fix the URL first
python3 - <<'PY'
import json
d = json.load(open("/tmp/models.json"))
ids = [m.get("id") for m in d.get("data", [])]
print(len(ids), "models returned")
print("\n".join(str(i) for i in ids[:10]))
PY
One more thing that command catches: if the first 200 bytes are HTML rather than JSON, something between you and the endpoint answered — a captive portal, a corporate proxy — and no amount of config change will fix it from inside your app. That is the same first check as in base URL not working, and it is worth running before the model-name work rather than after.
What resolving to the wrong model costs
A name that resolves silently is worse than one that errors. If a fallback or an alias points somewhere you did not intend, every call succeeds, nothing raises, and the invoice reflects a model you never chose.
| Model | Gateway rate in / out per 1M tokens | Official list in / out per 1M tokens | Diff |
|---|---|---|---|
| gpt-5.6-luna | $0.1 / $0.6 | $0.2 / $1.2 | 50% |
| deepseek-v4-flash | $0.22 / $0.66 | — / — | — |
| claude-sonnet-5 | $1 / $5 | $2 / $10 | 50% |
Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.
Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget. Across the 216 models in our cleaned catalogue the median output-to-input ratio is 4.0×.
Related guides
- One endpoint for every tool — the hub for tool-specific setup
- LiteLLM with a custom endpoint — getting the base URL right first
- LiteLLM retries and timeouts — the budget after names resolve
- Endpoint error codes: 401 vs 404 — routing vs credential vs path
- Reading /v1/models — where the authoritative id comes from
- Base URL not working — when the URL is the problem, not the name
Name-resolution failures and what each one is actually telling you
| Symptom | What is actually happening | How to confirm | Fix |
|---|---|---|---|
| BadRequestError: LLM Provider NOT provided | No provider in the string and no match in the internal map; raised before the request is sent. | Run the split function above on the exact string. | Prefix with openai/, or pass custom_llm_provider explicitly. |
| 404 on a model id you copied from documentation | The endpoint registers a different id than the docs show — often with a vendor segment. | Compare against the endpoint's /v1/models output. | Send the registered id exactly; stack the prefix if it contains a slash. |
| A request with no prefix worked yesterday and fails today | Provider inference depends on LiteLLM's internal model map, which changes between versions. | Pin your LiteLLM version and re-run the split check. | Write the prefix explicitly so inference is never involved. |
| Changing max_retries or num_retries changes nothing about the error | It is a client-side routing error, so no retry loop is entered. | Check whether the error is raised before any HTTP call. | Fix the model string; leave the retry policy alone. |
| The proxy answers but the upstream id in the logs is unfamiliar | An alias in model_list points somewhere other than you think, or a fallback absorbed the call. | Print the resolved litellm_params.model for the alias you called. | Audit model_list so every alias names exactly one qualified id. |
| Costs change with no code change | An upstream alias or default model was updated underneath a stable model_name. | Plot cost per request by resolved model, not by alias. | Pin aliases to explicit ids and review them like any other dependency. |
When this is the wrong move
Four shortcuts that turn a two-minute fix into an afternoon
| Situation | Why it breaks | Do this instead |
|---|---|---|
| You are about to add retries to fix LLM Provider NOT provided | The error is raised before a request exists, so every retry fails identically and nothing is billed either way. | Fix the string, then measure. |
| You are about to rotate a key because a model name 404s | A 404 means routing and authentication both succeeded; the credential is proven good by the response. | Read the id out of /v1/models and resend. |
| You are copying a model id out of a URL slug | Slugs are rewritten for the address bar; ids are what the API compares against. | Copy from the /v1/models response, or from a rate table that quotes the id. |
| You are relying on provider inference in production | Inference depends on a model map that ships with the library and changes between releases. | Write the prefix, and pin the version. |
Rolling it out without finding out the hard way
The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:
- Prove it on one request per alias, with the resolved model id printed back. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
- Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
- Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
- Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.
Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.
FAQ
What does the <code>openai/</code> prefix actually do?
It is a routing instruction, not part of the model's name. LiteLLM's documentation for OpenAI-compatible endpoints is explicit: put openai/ in front of the model name so LiteLLM knows you are calling a /chat/completions endpoint, and LiteLLM then strips that first segment and sends the remainder as the model value. Because only the first slash is treated as the separator, anything after it reaches the endpoint untouched — which is why a model id that itself contains a slash has to be stacked, as in openai/vendor/model.
What does <code>LLM Provider NOT provided</code> mean, and did that request cost anything?
It means LiteLLM could not work out which provider adapter to use, and it is raised before anything is sent. The full message names the string you passed: BadRequestError: LLM Provider NOT provided. Pass in the LLM provider you are trying to call. You passed model=support-bot. No HTTP request leaves the process, so there is no upstream call and nothing is billed. That is the useful half of the diagnosis — it separates a routing problem, which is free, from a 404, which means routing succeeded and the id was wrong.
Why does my model id 404 even though the prefix is correct?
Usually because the id you are sending is not the id the endpoint registered. LiteLLM passes the post-slash remainder through unchanged, so if the endpoint's own list shows vendor/model and you send openai/model, the endpoint receives model and rejects it. Read the registered id out of GET /v1/models and send exactly that string. The second common cause is a base URL with the wrong suffix — LiteLLM's docs note that a missing /v1 produces a not-found error, while a doubled one produces /v1/v1/chat/completions.
Can I avoid putting the prefix in the model string?
Yes, two ways. Pass custom_llm_provider="openai" as a separate argument when the string itself comes from somewhere you cannot edit — a database column, a framework config, a UI field. Or, if you run the LiteLLM proxy, keep client code clean and map a friendly model_name to a fully qualified litellm_params.model in the config. The second is the better long-term shape: the backend becomes a one-line change.
Does the prefix change the price?
No. The prefix decides which adapter formats the request; it has no effect on the tokens sent or on the rate charged. What changes the price is which model the id resolves to, which is exactly why the mapping is worth keeping somewhere you can audit. Our rate tables are the reference for what each id currently costs through the gateway.
Do I need a prefix if I only ever call one endpoint?
You need a provider decision, and the prefix is the cheapest way to make it explicit. LiteLLM can often infer the provider for well-known OpenAI or Anthropic ids, but inference depends on the internal model map in the version you have installed; a bare id that stops being recognised after an upgrade produces the routing error above. Writing the prefix costs one line and removes that dependency entirely.