Cherry Studio with a Custom OpenAI-Compatible Provider
Cherry Studio is a desktop chat client with a provider list, a key manager and a model picker. Pointing it at your own OpenAI-shaped endpoint takes about a minute, and the minute is where the mistakes happen — because the address field is not a base URL field, the model list is not populated from your server, and the provider is created in a disabled state.
The API Address field appends /v1/chat/completions to whatever you type. Paste a complete endpoint and you get the path twice. A trailing # is Cherry Studio's own switch for turning that append off — and it is the only client on this site that solves the problem with a character in the address. Models are never discovered for you: the picker shows the list you added, not the list your gateway serves. And the Enable toggle is off on new providers, so a correct configuration looks like nothing happened.
https://aicomp.ai/v1).
Create one free →
This page covers the exact field order in the GUI, what the address field does to your input, how the model list really works, and what a month of desktop chat costs once conversation history is accounted for.
The path through the UI
- Settings → Model Services, then the add button to create a provider.
- Choose the provider type.
OpenAIis the one that sends an OpenAI-shaped request. Cherry Studio also ships aNew APItype that speaks several protocols from one address — see below. - Name it anything you will recognise later. The name is only a label.
- Paste the API key. Multiple keys go in as a comma-separated list, in ASCII commas.
- Enter the API Address — the root address of the gateway, not the completions URL.
- Fetch or add models, then click Test with a model selected.
- Switch the provider on with the toggle in the top-right of the provider page.
Step 7 is the one people skip. The provider exists, the fields are filled, and the model dropdown stays empty until the toggle is on.
The address field appends the version path
The documented behaviour is that if a provider hands you a full URL such as https://xxx.com/v1/chat/completions, you enter the base URL https://xxx.com, and Cherry Studio appends /v1/chat/completions itself.
So the field wants a prefix, and the failure mode of giving it more than that is a doubled path rather than a clear error:
https://gw.example.com/v1 → https://gw.example.com/v1/v1/chat/completions
For a gateway that does not serve the standard path, the address field takes the full address followed by #. Per the documentation, an address ending in # stops the version path from being appended and is used exactly as entered. Two things follow from that:
- The
#is a terminator, not part of the URL. It never leaves the client, and it must be the last character. - You should reach for it only when the gateway genuinely needs a non-standard path. For a normal OpenAI-compatible deployment the root address is correct, and adding
#because requests are failing moves the problem somewhere harder to see.
Standard path → https://gw.example.com
Non-standard path → https://gw.example.com/openai/v1/chat/completions#
If you are running a multi-protocol proxy, the same rule applies per protocol: a previous beta configuration that left localhost:3000 in one of the secondary address fields will keep sending those models to the old host. Clear the stale address before testing.
Models are added, not discovered
The model list is a local list with a fetch helper attached, and the distinction matters:
- Only models present in the list appear in Cherry Studio's model selectors. A gateway serving forty models shows zero until you add some.
- "Get model list" queries the endpoint and opens a management list where you add the ones you want. Adding manually is equally valid and is what you do when the endpoint does not implement the models route.
- The real API model ID is displayed under each name, and near-duplicates — date-suffixed revisions, vendor-prefixed variants — are deliberately not merged. Pick by the ID, not by the label, and use the ID your endpoint actually accepts.
This is the same asymmetry that appears in every client on this site: the display name is yours, the model field is the endpoint's, and only the second one is sent.
The key field accepts several keys
Cherry Studio supports multiple API keys for one provider and rotates through them from front to back. Quick entry is a comma-separated list — and the commas must be ASCII. A full-width comma looks identical in most editors, produces a single nonsense key, and returns the same 401 as a wrong key. The alternative is the key manager, where each key is a separate entry with its own label, enabled state and delete action, which is what you want for a shared machine.
When one address should speak several protocols
The built-in New API provider type connects OpenAI Chat, OpenAI Responses, Anthropic Messages and Gemini protocols from a single address at the same time, choosing the versioned path per protocol. If your gateway is a multi-protocol proxy, this is a shorter configuration than creating one provider per protocol — and the same address rule applies: give the root address and let Cherry Studio pick /v1 or /v1beta, and only append # if the proxy explicitly requires a path with no version.
There is also an advanced API Settings panel per provider for capability flags — array-form message content, developer messages, reasoning toggles for models that expose them, and provider-specific settings such as service_tier — which you normally leave alone.
Why the input column dominates here
A desktop chat client re-sends the whole conversation on every turn. Nothing is cached between messages and nothing is summarised on your behalf, so a long thread with a large file attached pays for that file again on every subsequent question. Output is one reply per turn.
A day of work on this page's workload is 170k input and 30k output tokens — 3.4M input and 0.6M output across 20 working days. Because the input side grows with the thread rather than with the answer, the input rate is the number that decides the month: see why output pricing dominates the bill for the other half of the equation.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-sonnet-5 | $1.00 / $5.00 | $0.32 | $6.40 |
| gpt-5.6-luna | $0.10 / $0.60 | $0.04 | $0.70 |
| deepseek-v4-flash | $0.22 / $0.66 | $0.06 | $1.14 |
| MiniMax-M3 | $0.15 / $0.60 | $0.04 | $0.87 |
| qwen3.8-max | $1.00 / $3.00 | $0.26 | $5.20 |
One day = 170k input + 30k output on this page's workload. 20 days = 3.4M input and 0.6M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
Verifying the route
- Click Test with a model selected. A successful test proves the address, the key and the model ID together — the three fields that fail independently everywhere else.
- Confirm the provider toggle is on if the model selector is empty. The test can pass while the provider is still disabled.
- Send one message in a new topic. A reply confirms the whole chain end to end; an empty reply with no error points at the address field, and a 401 points at the key.
- Read the request path in the gateway log. One
/v1is correct; two means you pasted a versioned URL into the address field; a path that does not match your deployment means the#terminator belongs there. - Add a second model from the same gateway and answer with it. That proves the list entries map to real model IDs rather than display names.
FAQ
FAQ
Why is my model list empty after a successful test?
Models are not imported automatically. Use the fetch action on the provider page to see what the endpoint serves, then add the ones you want — only added models appear in the pickers, and the provider's enable toggle has to be on as well.
Should the API Address end with `/v1`?
The field appends /v1/chat/completions itself, so the root address such as https://gw.example.com is what it expects. A /v1-suffixed address produces a doubled version segment.
What does the `#` at the end of the address do?
It disables Cherry Studio's automatic version-path append, so the address is used exactly as written. Use it only for gateways that serve chat completions at a non-standard path, and keep it as the final character.
Why does one of my two keys always fail?
The comma between them. Cherry Studio splits the key field on ASCII commas; a full-width comma is treated as part of the key, and the resulting credential produces a plain 401.
Can I use one provider for several protocols?
Yes — the New API provider type covers OpenAI Chat, OpenAI Responses, Anthropic Messages and Gemini from a single address, choosing the versioned path per protocol.
Is Cherry Studio cheaper per prompt than a CLI agent?
Lower per prompt, because a chat turn writes no files. Higher per long thread, because nothing is trimmed. For the opposite shape — high absolute input driven by tool results — see Crush, and for the same conversation-replay pattern with environment-variable configuration, see LobeChat.
Related
- LobeChat — the self-hosted chat client, configured by
OPENAI_PROXY_URLrather than a GUI - Open WebUI — the browser-based alternative, with its own endpoint field
- AnythingLLM — same provider-and-model-list pattern, aimed at document workspaces
- Tools that take one endpoint — the full matrix of clients with a custom base URL field
- Official list prices vs gateway rates — which of the two numbers on this page is which