Cursor with an OpenAI-Compatible Endpoint: Base URL and Custom Models
You opened Cursor Settings → Models, switched on Override OpenAI Base URL, pasted an endpoint, put a key in the OpenAI API Key field, and pressed Verify. Cursor said the connection worked. Then you opened the model picker and every model you actually wanted to use was missing — or the one you typed in came back model not found.
Both outcomes are normal, and neither means the endpoint is broken. They come from two things about how this setting works that the field label does not tell you.
First, Verify only proves that the host answered an authenticated request. It does not prove the model you intend to use exists there. Second, Cursor does not read the model list from your endpoint. There is no step where it calls /v1/models and populates the picker. "Refresh Model List" refreshes Cursor's own built-in catalogue, not anything served by the URL you just entered. Models have to be added by hand, with the exact ID your endpoint expects.
This page covers the whole path: what the override actually replaces, the four URL formats that fail and why, adding models correctly, testing an endpoint before you trust it, and the failure modes that look like Cursor bugs but are not.
What the override actually does
By default, when you supply your own OpenAI key, Cursor sends OpenAI-shaped requests to OpenAI. The override replaces the destination. Cursor still builds the same request body and still appends the same path — it just sends it somewhere else.
That has three consequences worth internalising before you configure anything:
It only affects OpenAI-shaped traffic. The request is a POST to /chat/completions. An endpoint that speaks the OpenAI chat completions protocol will accept it; one that only implements the native Anthropic Messages protocol will not. If you want to use Claude models with your own key, that is the separate Anthropic field, not this one — see pointing Cursor at a custom Anthropic base URL for that path.
It does not touch Tab. Cursor's Tab completion runs on Cursor's own backend and is unaffected by any key you add. If you configured the override expecting autocomplete to start billing your key, it will not, and no setting will make it.
It does not touch subscription models. Models covered by your Cursor plan keep going through Cursor. The override applies to traffic that would otherwise have gone to OpenAI with your key.
The base URL: four formats that fail
Cursor appends the rest of the path itself. This is where most setups go wrong.
| Value you enter | Result |
|---|---|
https://example.ai/v1 | ✅ Correct |
https://example.ai/v1/chat/completions | ❌ Cursor appends the path again |
https://example.ai/v1/ | ❌ Trailing slash — Verify fails |
https://example.ai | ❌ Missing the version segment |
The second one is the intuitive mistake: you paste the full URL you saw in a snippet, and the request ends up at /v1/chat/completions/chat/completions. The third is the one that costs an afternoon, because it looks identical to the correct value and the only symptom is that Verify refuses to pass. A trailing slash is not trimmed.
The fourth fails for a different reason worth understanding, because it is the mirror image of a rule you may have read elsewhere. Native Anthropic clients — Claude Code, the Anthropic SDK, ChatAnthropic — append /v1/messages themselves, so they want a base with no /v1. Cursor is not one of them. It sends /chat/completions, and it wants the versioned base. The rule is not "always strip /v1" or "always add /v1"; it is whatever your client does not add itself.
https://aicomp.ai/v1).
Create one free →
Adding the models by hand
This is the step most guides skip, and it is the reason a verified endpoint still produces a working setup with no usable models.
After the override is saved, open Models and use + Add Model. Type the identifier exactly as your endpoint publishes it. deepseek-v4-flash is not deepseek-flash, and a name that is off by one character returns model not found on the first message even though Verify passed.
To find the exact IDs, ask the endpoint rather than guessing:
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
The response is a JSON list of id values. Copy them character for character. Partial IDs, aliases, and the names you are used to from another provider are the three most common sources of model not found here.
A related trap: adding the model makes it appear in the picker, but appearing is not the same as working. Send one real message before you assume the wiring is done.
Verify proves less than it looks like
The Verify button checks that the host is reachable and that the key is accepted. It is a connectivity and authentication test. It does not exercise streaming, tool calling, or long contexts — the paths a real agent session actually uses.
So a green Verify followed by a broken first message is a normal sequence, and it usually means one of:
- the endpoint does not implement streaming the way Cursor expects it
- the model does not support tool calls, and the agent path needs them
- the model ID resolves but the context window is smaller than the request Cursor built
Test end to end. One message in a fresh chat, then check your gateway's usage log for today's date. A request appearing there is the only proof that traffic is actually billing to your key.
Test the endpoint before you paste it
There is a cheap way to tell whether a URL genuinely implements the OpenAI protocol, and it works even with no usable key. Send a request with a deliberately wrong key:
curl -s https://example.ai/v1/chat/completions \
-H "Authorization: Bearer sk-deliberately-invalid" \
-H "Content-Type: application/json" \
-d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"hi"}]}'
Read the shape of the response, not the status code:
- A JSON auth error — something like
Invalid tokenorinvalid_api_key— means the route exists and is real. The endpoint is genuine; you just used a bad key. - An HTML page — a marketing homepage or a generic 404 document — means this path was never implemented. No key will make it work, and pasting it into Cursor will fail in a way that looks like a Cursor problem.
This takes ten seconds and saves the entire "is it me or is it them" loop.
What a Cursor session actually costs
Cursor is not a chatbot. Every message carries repository context, so input dominates the bill and the numbers do not resemble what you would pay for the same model in a chat window. A single agent task can read a dozen files before it writes anything, and an agent loop re-reads whatever is in scope on each turn. A realistic day lands near 150k input tokens and 40k output tokens. Over 20 working days that is 3M input and 0.8M output.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-sonnet-5 | $1.00 / $5.00 | $0.35 | $7.00 |
| gpt-5.6-sol | $2.5 / $15.00 | $0.97 | $19.50 |
| deepseek-v4-flash | $0.22 / $0.66 | $0.06 | $1.19 |
| glm-5.3 | $0.70 / $2.2 | $0.19 | $3.86 |
| qwen3.8-max | $1.00 / $3.00 | $0.27 | $5.40 |
One day = 150k input + 40k output on this page's workload. 20 days = 3M input and 0.8M output. Rates checked 2026-09-26 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
Two things follow from that table. The spread between the most and least expensive option is large enough that model choice is the single biggest lever you have — bigger than any prompt-tuning trick. And because input dominates, the models that look cheap per output token are not automatically the cheap ones for Cursor; check the input column first.
One structural fact about model pricing that surprises people: output is priced above input on essentially every model (
). Cursor inverts the usual weighting by sending so much context, which is exactly why reading the input column matters more here than it does anywhere else in your stack.
When the override is the wrong tool
Do not use it if you only want Claude models on your own key. That is the Anthropic field, and mixing the two is a known source of confusing behaviour — see Claude API in Cursor for the case where the label says Anthropic but the wire format is not.
Do not use it if you rely on Tab. That path is unaffected.
And if requests are failing in ways that look like routing problems rather than key problems, base URL not working separates the two families of failure by the status code you get back.
FAQ
FAQ
Does the override affect Cursor Tab?
No. Tab completion runs on Cursor's own backend regardless of any key or base URL you configure. Only chat, Composer and inline edit can route through your endpoint.
Why did Verify pass but the first message fail?
Verify tests connectivity and authentication only. It does not exercise streaming, tool calls or long contexts. A model that authenticates fine can still fail on any of those three.
Why does my model not appear after saving the base URL?
Because Cursor does not enumerate custom endpoints. Use + Add Model and type the ID exactly as your endpoint publishes it. "Refresh Model List" only refreshes Cursor's built-in catalogue.
Should the base URL end in /v1?
Yes, for Cursor. It sends /chat/completions and expects the versioned base. The "no /v1" rule applies to native Anthropic clients such as Claude Code and the Anthropic SDK, which append /v1/messages themselves.
Does a trailing slash matter?
Yes. https://example.ai/v1/ fails verification while https://example.ai/v1 passes. The slash is not trimmed.
Can I use one key across several providers?
Only if the endpoint you are pointing at aggregates them. A single vendor's endpoint will reject model IDs belonging to another vendor regardless of which key you use.
Related
- Cursor with a custom Anthropic base URL — the other override, and which one to use when
- Claude API in Cursor — where the Anthropic label does not mean the Anthropic wire format
- Cursor API key errors — 401s, 403s and key scoping in Cursor specifically
- Base URL not working — separate routing failures from credential failures by status code
- All OpenAI-compatible tooling — the same pattern across 30+ editors, CLIs and SDKs
- Full rate index — every model with a published rate