LobeChat with a Custom OpenAI-Compatible Endpoint

In short: The self-hosted variable is OPENAI_PROXY_URL, not the OPENAI_BASE_URL most clients read — anything else is ignored silently and requests keep going to the default host with a key that cannot work there. Whether the URL needs a /v1 suffix depends on your gateway; the documented symptom of getting it wrong is an empty reply, not a 404. Visible models are controlled by OPENAI_MODEL_LIST, not discovered.

LobeChat is an open-source chat UI — web, desktop and Docker — that talks to any OpenAI-compatible endpoint rather than to one vendor. It is the tool people reach for when they want one interface over several models, which is exactly the case where pointing it at a gateway pays off.

There are two configuration routes, and three things about them are easy to get wrong. The self-hosted variable is OPENAI_PROXY_URL, not the OPENAI_BASE_URL most clients use — set the wrong one and LobeChat keeps calling the real OpenAI, now with a key that cannot work there. Whether the URL needs a /v1 suffix depends on your gateway, and LobeChat's own documentation says to try adding it when replies come back empty — the symptom is a blank response, not the 404 you would expect from a wrong path. And models do not appear just because your gateway serves them: the visible list is controlled by OPENAI_MODEL_LIST, with its own +, - and -all syntax.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page covers both routes, the variable name, how to settle the /v1 question from evidence, the model list syntax that decides what you can actually pick, and what a month of chat traffic costs once you account for history being re-sent on every turn.

Two routes, and when to use each

Route one: the settings panel. Use this for the hosted web app or the desktop client. Settings → AI Service Provider → OpenAI, then:

FieldValue
API Keyyour gateway key
API Proxy URLhttps://your-gateway.example.com/v1

Enable the custom-endpoint toggle if your build shows one — without it the proxy field is ignored and requests go to the default host.

Route two: environment variables. Use this for Docker, docker-compose, Vercel or Kubernetes.

docker run -d --name lobe-chat -p 3210:3210 \
  -e OPENAI_API_KEY=your-key-here \
  -e OPENAI_PROXY_URL=https://your-gateway.example.com/v1 \
  -e OPENAI_MODEL_LIST="-all,+claude-sonnet-5,+gpt-5.6-sol" \
  lobehub/lobe-chat

The variable is OPENAI_PROXY_URL

This is the one that catches people coming from another client:

VariableMeaning
OPENAI_API_KEYthe key sent as the bearer token
OPENAI_PROXY_URLoverrides the request base URL. Default https://api.openai.com/v1

It is not OPENAI_BASE_URL, OPENAI_API_BASE or OPENAI_ENDPOINT. Set one of those and nothing happens — no warning, no error about an unknown variable. LobeChat keeps its default host and sends your gateway key to the real OpenAI, where it is rejected. The failure looks like a bad key, so people re-generate the key instead of renaming the variable.

Whether the URL ends in /v1

LobeChat's own environment-variable documentation is unusually honest about this being unknowable in advance:

Please check the request suffix of your proxy service provider. Some proxy service providers may add /v1 to the request suffix, while others may not. If you find that the AI returns an empty message during testing, try adding the /v1 suffix and retry.

Whether to include it depends on whether the host in front of you forwards the versioned path already. The documentation points at several long-running issues about replies coming back blank, which tells you what the failure actually looks like:

Settle it from the request path your gateway recorded:

POST /v1/chat/completions        ← correct
POST /v1/v1/chat/completions     ← your URL ends in /v1 and the client adds one
POST /chat/completions           ← the host forwards the versioned path, so drop yours

The model list is not automatic

OPENAI_MODEL_LIST controls what you can pick. The syntax is its own small language:

SyntaxEffect
+model-idadd a model
-model-idhide a model
-alldisable every default first
model-id=Display Namerename in the picker
,separator — no spaces
# only your gateway's models, with friendlier names
OPENAI_MODEL_LIST="-all,+claude-sonnet-5=Claude Sonnet 5,+gpt-5.6-sol=GPT-5.6 Sol"

Three things follow from this:

Across 219 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (150 models) charge at least 4× more for output than input. The other 624 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Why the input column dominates here

A chat UI re-sends the conversation with every turn. Turn ten carries turns one through nine, so the input side grows with the length of the conversation while the output side stays roughly flat per message. Two LobeChat-specific habits make it worse: long threads that are never branched, and the multi-model comparison view, which asks the same question of several models and pays for each answer.

A day of work on this page's workload is 150k input and 30k output tokens — 3M input and 0.6M output across 20 working days. Compare models deliberately rather than continuously, and start a new thread when the topic changes — both cut the input column without touching what you ask for. See what actually drives agent cost for the same effect measured on a coding workload.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-sonnet-5$1.00 / $5.00$0.30$6.00
gpt-5.6-sol$2.5 / $15.00$0.82$16.50
deepseek-v4-flash$0.22 / $0.66$0.05$1.06
glm-5.3$0.70 / $2.2$0.17$3.42
kimi-k3$1.5 / $7.5$0.45$9.00

One day = 150k input + 30k output on this page's workload. 20 days = 3M input and 0.6M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Verifying the route

  1. Send a one-line message. A reply proves key, host, path and model id at once.
  2. An empty reply with no error: check /v1 before anything else.
  3. A 401: check the variable name. OPENAI_BASE_URL does nothing here.
  4. A model that is missing from the picker: OPENAI_MODEL_LIST, and remember -all comes first.
  5. A model that is listed but fails on send: the id does not match what your gateway serves.

FAQ

FAQ

Why does LobeChat still call api.openai.com?

The proxy variable is OPENAI_PROXY_URL. Any other name is ignored silently, and the default host stays in place.

Do I need `/v1` at the end?

It depends on whether your gateway forwards the version prefix. The official guidance is to add it and retry if replies come back empty; the gateway request log gives you the definitive answer.

Why is my model not in the picker?

OPENAI_MODEL_LIST. Add it with +, and use -all first if the default OpenAI list is in the way.

Can I rename models in the list?

Yes — model-id=Display Name. The id sent to the API is unaffected.

Do model ids with a slash work?

Yes. Ids are treated as strings and passed through, so provider/model forms are fine.

Is the multi-model comparison expensive?

It is N requests instead of one against the same growing history. Worth it when you are choosing a model, wasteful as a default.

Related

Get API access