Cursor with an OpenAI-Compatible Endpoint: Base URL and Custom Models

You opened Cursor Settings → Models, switched on Override OpenAI Base URL, pasted an endpoint, put a key in the OpenAI API Key field, and pressed Verify. Cursor said the connection worked. Then you opened the model picker and every model you actually wanted to use was missing — or the one you typed in came back model not found.

Both outcomes are normal, and neither means the endpoint is broken. They come from two things about how this setting works that the field label does not tell you.

First, Verify only proves that the host answered an authenticated request. It does not prove the model you intend to use exists there. Second, Cursor does not read the model list from your endpoint. There is no step where it calls /v1/models and populates the picker. "Refresh Model List" refreshes Cursor's own built-in catalogue, not anything served by the URL you just entered. Models have to be added by hand, with the exact ID your endpoint expects.

This page covers the whole path: what the override actually replaces, the four URL formats that fail and why, adding models correctly, testing an endpoint before you trust it, and the failure modes that look like Cursor bugs but are not.

What the override actually does

By default, when you supply your own OpenAI key, Cursor sends OpenAI-shaped requests to OpenAI. The override replaces the destination. Cursor still builds the same request body and still appends the same path — it just sends it somewhere else.

That has three consequences worth internalising before you configure anything:

It only affects OpenAI-shaped traffic. The request is a POST to /chat/completions. An endpoint that speaks the OpenAI chat completions protocol will accept it; one that only implements the native Anthropic Messages protocol will not. If you want to use Claude models with your own key, that is the separate Anthropic field, not this one — see pointing Cursor at a custom Anthropic base URL for that path.

It does not touch Tab. Cursor's Tab completion runs on Cursor's own backend and is unaffected by any key you add. If you configured the override expecting autocomplete to start billing your key, it will not, and no setting will make it.

It does not touch subscription models. Models covered by your Cursor plan keep going through Cursor. The override applies to traffic that would otherwise have gone to OpenAI with your key.

The base URL: four formats that fail

Cursor appends the rest of the path itself. This is where most setups go wrong.

Value you enterResult
https://example.ai/v1✅ Correct
https://example.ai/v1/chat/completions❌ Cursor appends the path again
https://example.ai/v1/❌ Trailing slash — Verify fails
https://example.ai❌ Missing the version segment

The second one is the intuitive mistake: you paste the full URL you saw in a snippet, and the request ends up at /v1/chat/completions/chat/completions. The third is the one that costs an afternoon, because it looks identical to the correct value and the only symptom is that Verify refuses to pass. A trailing slash is not trimmed.

The fourth fails for a different reason worth understanding, because it is the mirror image of a rule you may have read elsewhere. Native Anthropic clients — Claude Code, the Anthropic SDK, ChatAnthropic — append /v1/messages themselves, so they want a base with no /v1. Cursor is not one of them. It sends /chat/completions, and it wants the versioned base. The rule is not "always strip /v1" or "always add /v1"; it is whatever your client does not add itself.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

Adding the models by hand

This is the step most guides skip, and it is the reason a verified endpoint still produces a working setup with no usable models.

After the override is saved, open Models and use + Add Model. Type the identifier exactly as your endpoint publishes it. deepseek-v4-flash is not deepseek-flash, and a name that is off by one character returns model not found on the first message even though Verify passed.

To find the exact IDs, ask the endpoint rather than guessing:

curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

The response is a JSON list of id values. Copy them character for character. Partial IDs, aliases, and the names you are used to from another provider are the three most common sources of model not found here.

A related trap: adding the model makes it appear in the picker, but appearing is not the same as working. Send one real message before you assume the wiring is done.

Verify proves less than it looks like

The Verify button checks that the host is reachable and that the key is accepted. It is a connectivity and authentication test. It does not exercise streaming, tool calling, or long contexts — the paths a real agent session actually uses.

So a green Verify followed by a broken first message is a normal sequence, and it usually means one of:

Test end to end. One message in a fresh chat, then check your gateway's usage log for today's date. A request appearing there is the only proof that traffic is actually billing to your key.

Test the endpoint before you paste it

There is a cheap way to tell whether a URL genuinely implements the OpenAI protocol, and it works even with no usable key. Send a request with a deliberately wrong key:

curl -s https://example.ai/v1/chat/completions \
  -H "Authorization: Bearer sk-deliberately-invalid" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-5.6-sol","messages":[{"role":"user","content":"hi"}]}'

Read the shape of the response, not the status code:

This takes ten seconds and saves the entire "is it me or is it them" loop.

What a Cursor session actually costs

Cursor is not a chatbot. Every message carries repository context, so input dominates the bill and the numbers do not resemble what you would pay for the same model in a chat window. A single agent task can read a dozen files before it writes anything, and an agent loop re-reads whatever is in scope on each turn. A realistic day lands near 150k input tokens and 40k output tokens. Over 20 working days that is 3M input and 0.8M output.

Per 1M input / output tokens,
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-sonnet-5$1.00 / $5.00$0.35$7.00
gpt-5.6-sol$2.5 / $15.00$0.97$19.50
deepseek-v4-flash$0.22 / $0.66$0.06$1.19
glm-5.3$0.70 / $2.2$0.19$3.86
qwen3.8-max$1.00 / $3.00$0.27$5.40

One day = 150k input + 40k output on this page's workload. 20 days = 3M input and 0.8M output. Rates checked 2026-09-26 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Two things follow from that table. The spread between the most and least expensive option is large enough that model choice is the single biggest lever you have — bigger than any prompt-tuning trick. And because input dominates, the models that look cheap per output token are not automatically the cheap ones for Cursor; check the input column first.

One structural fact about model pricing that surprises people: output is priced above input on essentially every model (

Across 216 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (147 models) charge at least 4× more for output than input. The other 623 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

). Cursor inverts the usual weighting by sending so much context, which is exactly why reading the input column matters more here than it does anywhere else in your stack.

When the override is the wrong tool

Do not use it if you only want Claude models on your own key. That is the Anthropic field, and mixing the two is a known source of confusing behaviour — see Claude API in Cursor for the case where the label says Anthropic but the wire format is not.

Do not use it if you rely on Tab. That path is unaffected.

And if requests are failing in ways that look like routing problems rather than key problems, base URL not working separates the two families of failure by the status code you get back.

FAQ

FAQ

Does the override affect Cursor Tab?

No. Tab completion runs on Cursor's own backend regardless of any key or base URL you configure. Only chat, Composer and inline edit can route through your endpoint.

Why did Verify pass but the first message fail?

Verify tests connectivity and authentication only. It does not exercise streaming, tool calls or long contexts. A model that authenticates fine can still fail on any of those three.

Why does my model not appear after saving the base URL?

Because Cursor does not enumerate custom endpoints. Use + Add Model and type the ID exactly as your endpoint publishes it. "Refresh Model List" only refreshes Cursor's built-in catalogue.

Should the base URL end in /v1?

Yes, for Cursor. It sends /chat/completions and expects the versioned base. The "no /v1" rule applies to native Anthropic clients such as Claude Code and the Anthropic SDK, which append /v1/messages themselves.

Does a trailing slash matter?

Yes. https://example.ai/v1/ fails verification while https://example.ai/v1 passes. The slash is not trimmed.

Can I use one key across several providers?

Only if the endpoint you are pointing at aggregates them. A single vendor's endpoint will reject model IDs belonging to another vendor regardless of which key you use.

Related

Get API access

Get API access