LobeChat with a Custom OpenAI-Compatible Endpoint
LobeChat is an open-source chat UI — web, desktop and Docker — that talks to any OpenAI-compatible endpoint rather than to one vendor. It is the tool people reach for when they want one interface over several models, which is exactly the case where pointing it at a gateway pays off.
There are two configuration routes, and three things about them are easy to get wrong. The self-hosted variable is OPENAI_PROXY_URL, not the OPENAI_BASE_URL most clients use — set the wrong one and LobeChat keeps calling the real OpenAI, now with a key that cannot work there. Whether the URL needs a /v1 suffix depends on your gateway, and LobeChat's own documentation says to try adding it when replies come back empty — the symptom is a blank response, not the 404 you would expect from a wrong path. And models do not appear just because your gateway serves them: the visible list is controlled by OPENAI_MODEL_LIST, with its own +, - and -all syntax.
https://aicomp.ai/v1).
Create one free →
This page covers both routes, the variable name, how to settle the /v1 question from evidence, the model list syntax that decides what you can actually pick, and what a month of chat traffic costs once you account for history being re-sent on every turn.
Two routes, and when to use each
Route one: the settings panel. Use this for the hosted web app or the desktop client. Settings → AI Service Provider → OpenAI, then:
| Field | Value |
|---|---|
| API Key | your gateway key |
| API Proxy URL | https://your-gateway.example.com/v1 |
Enable the custom-endpoint toggle if your build shows one — without it the proxy field is ignored and requests go to the default host.
Route two: environment variables. Use this for Docker, docker-compose, Vercel or Kubernetes.
docker run -d --name lobe-chat -p 3210:3210 \
-e OPENAI_API_KEY=your-key-here \
-e OPENAI_PROXY_URL=https://your-gateway.example.com/v1 \
-e OPENAI_MODEL_LIST="-all,+claude-sonnet-5,+gpt-5.6-sol" \
lobehub/lobe-chat
The variable is OPENAI_PROXY_URL
This is the one that catches people coming from another client:
| Variable | Meaning |
|---|---|
OPENAI_API_KEY | the key sent as the bearer token |
OPENAI_PROXY_URL | overrides the request base URL. Default https://api.openai.com/v1 |
It is not OPENAI_BASE_URL, OPENAI_API_BASE or OPENAI_ENDPOINT. Set one of those and nothing happens — no warning, no error about an unknown variable. LobeChat keeps its default host and sends your gateway key to the real OpenAI, where it is rejected. The failure looks like a bad key, so people re-generate the key instead of renaming the variable.
Whether the URL ends in /v1
LobeChat's own environment-variable documentation is unusually honest about this being unknowable in advance:
Please check the request suffix of your proxy service provider. Some proxy service providers may add
/v1to the request suffix, while others may not. If you find that the AI returns an empty message during testing, try adding the/v1suffix and retry.
Whether to include it depends on whether the host in front of you forwards the versioned path already. The documentation points at several long-running issues about replies coming back blank, which tells you what the failure actually looks like:
- Empty or blank replies — the classic
/v1mismatch, not a model problem. - 404 — the path is wrong in the other direction, or the host does not serve chat completions at all.
Settle it from the request path your gateway recorded:
POST /v1/chat/completions ← correct
POST /v1/v1/chat/completions ← your URL ends in /v1 and the client adds one
POST /chat/completions ← the host forwards the versioned path, so drop yours
The model list is not automatic
OPENAI_MODEL_LIST controls what you can pick. The syntax is its own small language:
| Syntax | Effect |
|---|---|
+model-id | add a model |
-model-id | hide a model |
-all | disable every default first |
model-id=Display Name | rename in the picker |
, | separator — no spaces |
# only your gateway's models, with friendlier names
OPENAI_MODEL_LIST="-all,+claude-sonnet-5=Claude Sonnet 5,+gpt-5.6-sol=GPT-5.6 Sol"
Three things follow from this:
-allfirst. Without it you inherit the default OpenAI list, which your gateway does not serve — a picker full of models that fail on send.- Model ids are passed through as strings, so ids containing a slash work. See model name mapping for why the slash is part of the id and not a path.
- A model your gateway serves but you did not list is simply not there. Nothing is fetched from the endpoint to discover it.
Why the input column dominates here
A chat UI re-sends the conversation with every turn. Turn ten carries turns one through nine, so the input side grows with the length of the conversation while the output side stays roughly flat per message. Two LobeChat-specific habits make it worse: long threads that are never branched, and the multi-model comparison view, which asks the same question of several models and pays for each answer.
A day of work on this page's workload is 150k input and 30k output tokens — 3M input and 0.6M output across 20 working days. Compare models deliberately rather than continuously, and start a new thread when the topic changes — both cut the input column without touching what you ask for. See what actually drives agent cost for the same effect measured on a coding workload.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-sonnet-5 | $1.00 / $5.00 | $0.30 | $6.00 |
| gpt-5.6-sol | $2.5 / $15.00 | $0.82 | $16.50 |
| deepseek-v4-flash | $0.22 / $0.66 | $0.05 | $1.06 |
| glm-5.3 | $0.70 / $2.2 | $0.17 | $3.42 |
| kimi-k3 | $1.5 / $7.5 | $0.45 | $9.00 |
One day = 150k input + 30k output on this page's workload. 20 days = 3M input and 0.6M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
Verifying the route
- Send a one-line message. A reply proves key, host, path and model id at once.
- An empty reply with no error: check
/v1before anything else. - A 401: check the variable name.
OPENAI_BASE_URLdoes nothing here. - A model that is missing from the picker:
OPENAI_MODEL_LIST, and remember-allcomes first. - A model that is listed but fails on send: the id does not match what your gateway serves.
FAQ
FAQ
Why does LobeChat still call api.openai.com?
The proxy variable is OPENAI_PROXY_URL. Any other name is ignored silently, and the default host stays in place.
Do I need `/v1` at the end?
It depends on whether your gateway forwards the version prefix. The official guidance is to add it and retry if replies come back empty; the gateway request log gives you the definitive answer.
Why is my model not in the picker?
OPENAI_MODEL_LIST. Add it with +, and use -all first if the default OpenAI list is in the way.
Can I rename models in the list?
Yes — model-id=Display Name. The id sent to the API is unaffected.
Do model ids with a slash work?
Yes. Ids are treated as strings and passed through, so provider/model forms are fine.
Is the multi-model comparison expensive?
It is N requests instead of one against the same growing history. Worth it when you are choosing a model, wasteful as a default.
Related
- Open WebUI with a custom endpoint — the other major self-hosted chat UI, configured differently
- AnythingLLM — where embedding cost and query cost pull in opposite directions
- Output token pricing — why the output multiple, not the input rate, decides a chat bill
- Cheapest models by workload — pick by shape, not by headline rate