LiteLLM model names: the prefix is a routing instruction

LiteLLM reads your model string as two things, not one. It splits at the first slash: the segment before names the provider, and everything after is the model id handed to that provider untouched. Get that split wrong and you get one of two failures that look alike but are opposites — one raised before any request is sent, one returned by a server that received a perfectly routed call.

In short: LiteLLM splits the model string at the first slash: the first segment names the provider, the rest is passed upstream unchanged, so openai/vendor/model sends vendor/model. 'LLM Provider NOT provided' is raised before any request is sent, while a 404 means routing worked.
You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment
Confirm the key works first. One command, no SDK, costs nothing:
curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.

What the split actually does

LiteLLM's own documentation for OpenAI-compatible endpoints states the rule directly: put openai/ in front of your model name so LiteLLM knows you are calling a /chat/completions endpoint. For the older /completions route the documented prefix is text-completion-openai/ instead. The part after the first slash is passed to the provider unchanged.

"Unchanged" is the half people miss. It means additional slashes are not separators — they are part of the model id. So openai/glm/glm-5.2 resolves to provider openai and sends glm/glm-5.2 upstream. That is not a hack; it is the only way to reach an endpoint that registers its ids with a vendor segment already in them.

def split_model(model: str):
    """Mirror LiteLLM's documented rule: the first segment names the provider,
    everything after the first slash is the model id and is passed through
    to that provider unchanged -- including any further slashes."""
    provider, sep, rest = model.partition("/")
    if not sep:
        return None, model      # no provider in the string: routing must be inferred
    return provider, rest


for m in ["gpt-5.6-luna",
          "openai/gpt-5.6-luna",
          "openai/glm/glm-5.2",
          "anthropic/claude-sonnet-5",
          "support-bot"]:
    provider, sent = split_model(m)
    if provider is None:
        print(f"{m:30} -> no provider given, LiteLLM must guess or refuse")
    else:
        print(f"{m:30} -> provider={provider:10} sent upstream={sent}")

Run that against the ids you actually use before you debug anything else. It needs no LiteLLM installed, and it tells you in one line what the endpoint is going to receive.

The error that fires before the request does

Pass a string with no provider and no match in LiteLLM's internal map, and you get: litellm.exceptions.BadRequestError: LLM Provider NOT provided. Pass in the LLM provider you are trying to call. You passed model=support-bot.

Read the last clause literally — it is telling you the request never happened. No socket opened, no token counted, nothing billed. This is the single most useful fact on the page, because it splits the failure space cleanly:

Three ways to set the provider, and where each belongs

The prefix in the string is the most explicit and the easiest to read in a log. It is also the only one that survives being handed to a framework that re-serialises config.

import os
import litellm

# The prefix is a routing instruction, not part of the model name.
# LiteLLM strips the first segment and sends the rest to the endpoint as-is.
resp = litellm.completion(
    model="openai/gpt-5.6-luna",
    api_base="https://aicomp.ai/v1",
    api_key=os.environ["GATEWAY_API_KEY"],
    messages=[{"role": "user", "content": "Reply with the single word: ok"}],
    max_tokens=8,
)

print("answered by:", resp.model)      # identity, as the endpoint reports it
print("usage      :", resp.usage)      # None here means usage was never returned

# When you cannot change the model string (it comes from a database, a UI,
# a framework config), pass the provider as its own argument instead:
same = litellm.completion(
    model="gpt-5.6-luna",
    custom_llm_provider="openai",      # same routing, no prefix in the string
    api_base="https://aicomp.ai/v1",
    api_key=os.environ["GATEWAY_API_KEY"],
    messages=[{"role": "user", "content": "Reply with the single word: ok"}],
    max_tokens=8,
)

But the right place for the mapping, once you have more than one caller, is the proxy. There, model_name is what callers ask for and litellm_params.model is what is sent — so a friendly alias is decoupled from the qualified id behind it.

# litellm_config.yaml
# Callers ask for a friendly name; the proxy owns the provider-qualified one.
# Swapping the backend later is a one-line change here, not a client change.
model_list:
  - model_name: fast
    litellm_params:
      model: openai/deepseek-v4-flash-0731
      api_base: https://aicomp.ai/v1
      api_key: os.environ/GATEWAY_API_KEY
  - model_name: strong
    litellm_params:
      model: openai/claude-sonnet-5
      api_base: https://aicomp.ai/v1
      api_key: os.environ/GATEWAY_API_KEY

That indirection is what makes a backend swap a one-line change instead of a coordinated release across every service that hardcodes a model name. It is also the natural place to put a fallback, and therefore the place to think about how many upstream calls one request can turn into — which is why the retry and timeout budget is the next page to read once names resolve.

How many ids actually need the stacked form

In the 839-model catalogue we track, 12 ids contain a slash. For those, a single openai/ prefix is not enough — the vendor segment has to be preserved, which is exactly what stacking does.

Catalogue ids that carry their own vendor segment — the ones the stacked form exists for
Registered idWith the provider prefixSent upstream
BAAI/bge-reranker-v2-m3openai/BAAI/bge-reranker-v2-m3BAAI/bge-reranker-v2-m3
Pro/BAAI/bge-reranker-v2-m3openai/Pro/BAAI/bge-reranker-v2-m3Pro/BAAI/bge-reranker-v2-m3
Qwen/Qwen3-Reranker-0.6Bopenai/Qwen/Qwen3-Reranker-0.6BQwen/Qwen3-Reranker-0.6B
Qwen/Qwen3-Reranker-4Bopenai/Qwen/Qwen3-Reranker-4BQwen/Qwen3-Reranker-4B
Qwen/Qwen3-Reranker-8Bopenai/Qwen/Qwen3-Reranker-8BQwen/Qwen3-Reranker-8B
anthropic/claude-3.7-sonnetopenai/anthropic/claude-3.7-sonnetanthropic/claude-3.7-sonnet
anthropic/claude-sonnet-4openai/anthropic/claude-sonnet-4anthropic/claude-sonnet-4
meta-llama/llama-3.1-70b-instructopenai/meta-llama/llama-3.1-70b-instructmeta-llama/llama-3.1-70b-instruct

Showing 8 of 12. The remaining 827 ids carry no slash, so a single prefix is sufficient for them.

The general rule, and the only one worth memorising: never derive the id from anything except the endpoint's own list. Not from a docs page, not from a URL slug, not from what worked on a different endpoint.

# The endpoint's own list is the only authoritative source of model ids.
curl -s https://aicomp.ai/v1/models -H "Authorization: Bearer $GATEWAY_API_KEY" > /tmp/models.json

head -c 200 /tmp/models.json
#   -> starts with {"object":"list" ...   good
#   -> starts with <!DOCTYPE or <html     something else answered; fix the URL first

python3 - <<'PY'
import json
d = json.load(open("/tmp/models.json"))
ids = [m.get("id") for m in d.get("data", [])]
print(len(ids), "models returned")
print("\n".join(str(i) for i in ids[:10]))
PY

One more thing that command catches: if the first 200 bytes are HTML rather than JSON, something between you and the endpoint answered — a captive portal, a corporate proxy — and no amount of config change will fix it from inside your app. That is the same first check as in base URL not working, and it is worth running before the model-name work rather than after.

What resolving to the wrong model costs

A name that resolves silently is worse than one that errors. If a fallback or an alias points somewhere you did not intend, every call succeeds, nothing raises, and the invoice reflects a model you never chose.

Three ids a mistaken alias could land on — gateway rate vs official list
ModelGateway rate
in / out per 1M tokens
Official list
in / out per 1M tokens
Diff
gpt-5.6-luna$0.1 / $0.6$0.2 / $1.250%
deepseek-v4-flash$0.22 / $0.66— / ——
claude-sonnet-5$1 / $5$2 / $1050%

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget. Across the 216 models in our cleaned catalogue the median output-to-input ratio is 4.0×.

Cost note. The spread between the first and last row is the size of the mistake. A one-character difference in a model id can move a monthly bill by an order of magnitude, and because every request returns 200, nothing in your application will tell you it happened.

Related guides

Name-resolution failures and what each one is actually telling you

SymptomWhat is actually happeningHow to confirmFix
BadRequestError: LLM Provider NOT providedNo provider in the string and no match in the internal map; raised before the request is sent.Run the split function above on the exact string.Prefix with openai/, or pass custom_llm_provider explicitly.
404 on a model id you copied from documentationThe endpoint registers a different id than the docs show — often with a vendor segment.Compare against the endpoint's /v1/models output.Send the registered id exactly; stack the prefix if it contains a slash.
A request with no prefix worked yesterday and fails todayProvider inference depends on LiteLLM's internal model map, which changes between versions.Pin your LiteLLM version and re-run the split check.Write the prefix explicitly so inference is never involved.
Changing max_retries or num_retries changes nothing about the errorIt is a client-side routing error, so no retry loop is entered.Check whether the error is raised before any HTTP call.Fix the model string; leave the retry policy alone.
The proxy answers but the upstream id in the logs is unfamiliarAn alias in model_list points somewhere other than you think, or a fallback absorbed the call.Print the resolved litellm_params.model for the alias you called.Audit model_list so every alias names exactly one qualified id.
Costs change with no code changeAn upstream alias or default model was updated underneath a stable model_name.Plot cost per request by resolved model, not by alias.Pin aliases to explicit ids and review them like any other dependency.

When this is the wrong move

Four shortcuts that turn a two-minute fix into an afternoon

SituationWhy it breaksDo this instead
You are about to add retries to fix LLM Provider NOT providedThe error is raised before a request exists, so every retry fails identically and nothing is billed either way.Fix the string, then measure.
You are about to rotate a key because a model name 404sA 404 means routing and authentication both succeeded; the credential is proven good by the response.Read the id out of /v1/models and resend.
You are copying a model id out of a URL slugSlugs are rewritten for the address bar; ids are what the API compares against.Copy from the /v1/models response, or from a rate table that quotes the id.
You are relying on provider inference in productionInference depends on a model map that ships with the library and changes between releases.Write the prefix, and pin the version.

Rolling it out without finding out the hard way

The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:

  1. Prove it on one request per alias, with the resolved model id printed back. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
  2. Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
  3. Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
  4. Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.

Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.

FAQ

What does the <code>openai/</code> prefix actually do?

It is a routing instruction, not part of the model's name. LiteLLM's documentation for OpenAI-compatible endpoints is explicit: put openai/ in front of the model name so LiteLLM knows you are calling a /chat/completions endpoint, and LiteLLM then strips that first segment and sends the remainder as the model value. Because only the first slash is treated as the separator, anything after it reaches the endpoint untouched — which is why a model id that itself contains a slash has to be stacked, as in openai/vendor/model.

What does <code>LLM Provider NOT provided</code> mean, and did that request cost anything?

It means LiteLLM could not work out which provider adapter to use, and it is raised before anything is sent. The full message names the string you passed: BadRequestError: LLM Provider NOT provided. Pass in the LLM provider you are trying to call. You passed model=support-bot. No HTTP request leaves the process, so there is no upstream call and nothing is billed. That is the useful half of the diagnosis — it separates a routing problem, which is free, from a 404, which means routing succeeded and the id was wrong.

Why does my model id 404 even though the prefix is correct?

Usually because the id you are sending is not the id the endpoint registered. LiteLLM passes the post-slash remainder through unchanged, so if the endpoint's own list shows vendor/model and you send openai/model, the endpoint receives model and rejects it. Read the registered id out of GET /v1/models and send exactly that string. The second common cause is a base URL with the wrong suffix — LiteLLM's docs note that a missing /v1 produces a not-found error, while a doubled one produces /v1/v1/chat/completions.

Can I avoid putting the prefix in the model string?

Yes, two ways. Pass custom_llm_provider="openai" as a separate argument when the string itself comes from somewhere you cannot edit — a database column, a framework config, a UI field. Or, if you run the LiteLLM proxy, keep client code clean and map a friendly model_name to a fully qualified litellm_params.model in the config. The second is the better long-term shape: the backend becomes a one-line change.

Does the prefix change the price?

No. The prefix decides which adapter formats the request; it has no effect on the tokens sent or on the rate charged. What changes the price is which model the id resolves to, which is exactly why the mapping is worth keeping somewhere you can audit. Our rate tables are the reference for what each id currently costs through the gateway.

Do I need a prefix if I only ever call one endpoint?

You need a provider decision, and the prefix is the cheapest way to make it explicit. LiteLLM can often infer the provider for well-known OpenAI or Anthropic ids, but inference depends on the internal model map in the version you have installed; a bare id that stops being recognised after an upgrade produces the routing error above. Writing the prefix costs one line and removes that dependency entirely.

Get API access