Cursor Agent Mode with a Custom Endpoint: What Actually Routes

In short: The override changes where a class of requests goes, not every request Cursor makes. Chat routes through your endpoint; Tab completion always uses Cursor's own models; and background sub-agents have been observed ignoring the override entirely. Whether Agent mode routes depends on your Cursor build, so settle it from request logs, not tutorials.

You pointed Cursor at your own endpoint, got a reply in the chat panel, and assumed the expensive part was solved. Then you ran Composer on a real task and watched it behave in a way you cannot explain: no usage appeared on your provider dashboard for that request, or the agent stalled on the first tool call, or it worked Tuesday and stopped working after an update.

None of those are your imagination, and none of them mean the base URL is wrong. They come from something the Override OpenAI Base URL field does not tell you: enabling it does not make Cursor a thin client that forwards everything you do. Cursor's backend stays in the request path for prompt construction and context retrieval, and a handful of features never leave it at all. The override changes where a class of requests goes, and that class is smaller than the feature list in the UI suggests.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page maps that territory: which surfaces reach your endpoint, which are locked to Cursor's own models, where the boundary is documented versus merely observed, how to prove it from your side instead of trusting a forum thread, and what an Agent session costs once you know how many tokens it really sends.

What the override changes, precisely

Cursor builds a request and sends it somewhere. The override changes the somewhere, not the request.

That distinction carries most of the surprises. Cursor still assembles the messages array — system prompt, retrieved repository context, your file contents, conversation history — on its own servers. It still decides which path to append. It still decides, per feature, whether this request belongs to your credential or its own. So two things can be true at once: your key authenticates successfully, and a feature you expected to see on your bill never appears on it.

The requests that do honour the override are OpenAI-shaped. A POST to /chat/completions with a Bearer token, model and messages. An endpoint implementing that protocol accepts them; one speaking only the native Anthropic Messages protocol does not. If you want Claude models on your own key, that is the separate Anthropic field, not this one — see pointing Cursor at a custom Anthropic base URL for how those two fields differ, including the fact that Cursor sends OpenAI-shaped requests to both.

The boundaries Cursor documents

Two limits are stated plainly enough to treat as settled.

Chat models only. Cursor's BYOK documentation describes support for standard, non-reasoning chat models. "Available in the model picker" is the usable set, not "listed in your provider's catalogue". A gateway advertising 800 model IDs does not make 800 models reachable from your Cursor build.

Tab completion never routes. Tab runs on Cursor's own backend and models no matter what you configure. Seeing both Cursor subscription usage and provider usage is normal, not evidence of a broken setup.

The boundaries that are observed, not guaranteed

This is where forum threads disagree, and where it is worth being precise about what is actually known.

Several third-party integration guides state flatly that Cursor's coding agent — Composer, inline edit, apply/edit — does not route through external OpenAI-compatible endpoints, and that only the chat and plan panel honours a custom key. They tell you to use Claude Code or Cline if you need a fully routed agent. Other sources, including gateway operators documenting their own Cursor support, describe Agent and tool calls working, while listing release-specific payload bugs and version-gated fixes as ordinary hazards. One vendor records that image input did not reach custom endpoints until a specific release fixed the routing bug in June 2026.

Both can be true, because the answer depends on your Cursor build and release channel. What follows from that is not "pick a side" but stop relying on the claim and start testing locally. Three things you can establish in an afternoon:

Does traffic arrive at all? Open your provider or gateway dashboard and look at request logs rather than aggregate spend. If a Composer run produces no request with a matching timestamp, that feature did not use your credential. This is the single most useful measurement available to you, and it takes one run.

Does a request reach a second model family? Some gateway deployments expose a providerID/modelID path. If a nested model ID resolves while a bare one 404s, you have learned how your deployment addresses models, which is worth knowing before you debug anything else — the same ambiguity is covered in OpenCode's providerID/modelID addressing.

Does it survive an update? Whatever you conclude today, re-run the check after a Cursor update. Custom-endpoint behaviour has been patched repeatedly and silently.

The sub-agent exception

One boundary deserves its own warning, because it is the one that breaks governance rather than convenience.

When Cursor spawns a background sub-agent to work through a multi-step task, that agent has been observed falling back to Cursor's default routing and ignoring the configured base URL. This was reported and acknowledged as a known issue and remained open as of mid-2026.

The consequence is not a failed request — sub-agents generally succeed. The consequence is that requests you believed were covered by your gateway policy are not. If your organisation requires that every LLM call traverse an approved route for audit or data-residency reasons, Cursor cannot currently make that guarantee, and no amount of correct configuration will make it do so. Teams with that requirement are better served by a client whose entire call path is redirectable — ANTHROPIC_BASE_URL in Claude Code covers sub-agents and background agents too — and our comparison of routing control across coding agents goes into that asymmetry in detail.

Agent asks for more than a working chat request

A green Verify button and a paragraph of chat output prove exactly two things: the host answered, and a model produced text. Agent mode needs three more things from the same endpoint, and it needs them every turn.

Tool calls. The model has to emit structured tool calls and your endpoint has to return them intact. A model with weak or absent function-calling support answers chat beautifully and then stalls the moment a filesystem tool is needed. This is why "chat works, Agent hangs" is the single most reported shape of this problem, and why the advice to validate text first and tool behaviour second is worth following in that order.

Streaming. Cursor consumes Server-Sent Events. An endpoint that buffers, or a proxy that terminates idle connections mid-stream, produces a response that looks like a hang rather than an error. Corporate proxies interrupting HTTP/2 streams are common enough that Cursor exposes an explicit HTTP/1.1 compatibility mode — Settings, search for HTTP — and it is the first thing to try if Agent works at home and fails at the office.

Reasoning-field round-tripping. This one has a specific signature and it is worth knowing cold. Models that emit thinking content — several of the reasoning-class models including DeepSeek R1 derivatives and Kimi's thinking variants — expect that content echoed back in subsequent turns. A client that assembles the messages array without preserving it gets rejected on the second tool call, not the first: an error naming a missing reasoning field in an assistant tool call message. One symptom, one turn deep, and it only ever appears in Agent mode because plain chat never produces assistant messages with tool calls. The symptom is not unique to Cursor — it is a property of the model surfacing through whichever client drops the field, and the same reasoning-content signature appears elsewhere.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-sonnet-5$1.00 / $5.00$0.46$9.30
gpt-5.6-sol$2.5 / $15.00$1.27$25.50
deepseek-v4-flash$0.22 / $0.66$0.08$1.65
kimi-k3$1.5 / $7.5$0.70$13.95
glm-5.3$0.70 / $2.2$0.27$5.34

One day = 240k input + 45k output on this page's workload. 20 days = 4.8M input and 0.9M output. Rates checked 2026-09-26 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

What that table is showing you is the shape of the bill, and the shape is not what most people expect. A day of Agent work on this page's workload is 240k input and 45k output tokens — 4.8M input and 0.9M output across 20 working days. Every significant input cost is context Cursor re-sends on your behalf, not text you typed. Choosing a model by its output rate is the wrong reflex here; the input column decides the month.

Across 216 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (147 models) charge at least 4× more for output than input. The other 623 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Cursor inverts the usual weighting further than a terminal agent does, because each turn re-sends the retrieval results from the previous one. If you want the number rather than the principle, the Cursor cost calculator takes the workload and the model and does the arithmetic.

Proving which route a request took

When the question is genuinely "did that request use my key?", there is a reliable three-step order, cheapest test first.

One short message, no files, no tools. New chat, plain sentence. This separates "the endpoint is reachable" from "the endpoint supports what the feature needs". If this fails, nothing downstream matters — go to the general base URL failure list before assuming anything Cursor-specific.

One tool-using request. Anything that makes the model touch the filesystem. If text succeeds here and tools fail, you have localised it to model capability rather than connectivity, and the fix is a different model, not a different URL.

Then check the server, not the client. Read request logs filtered by timestamp, model ID and path. Local curl success establishes connectivity from your machine and nothing else — Cursor's requests originate from its servers, which is also why a localhost base URL cannot work and why a private IP is refused outright. Error shapes that look identical in the UI separate cleanly on the server by status code and key scope, and not at all in Cursor itself.

When Cursor is the wrong tool for this

Three situations where configuring harder is wasted effort.

You need every call routed. The sub-agent gap above. Use a client designed for it.

You only wanted to cut cost on autocomplete. Tab does not route. The spend you were targeting is not affected by anything you configure here.

You need a specific reasoning model in Agent mode. If the reasoning-field round trip above is the blocker, you are waiting on a client-side fix, not a configuration option. Chat-only use of that model is fine; Agent use is not.

None of these make the override a bad feature — it does exactly what it says for the surfaces it covers. They make it the wrong answer to three specific questions. Our guide to choosing a model by workload starts from the work rather than the tool, which usually produces a cheaper answer than optimising inside one editor.

FAQ

FAQ

Does Cursor Agent mode use my custom base URL?

It depends on your Cursor build and release channel, which is why integration guides disagree. Some state the coding agent never routes externally; others document it working with release-specific caveats. Establish it for your own setup by running one Composer task and checking whether a matching request appears in your endpoint's logs.

Why does my provider dashboard show less usage than I expected?

Because several Cursor surfaces do not use your credential. Tab completion always runs on Cursor's own models, and separately, sub-agents have been observed falling back to default routing. Low provider usage alongside working features is normal, not misconfiguration.

Can I route background sub-agents through my gateway?

Not reliably. The base URL being ignored for background agents was reported as a known issue and was still open as of mid-2026. If all traffic must traverse an approved route for audit or compliance reasons, treat Cursor as unable to guarantee it.

Why did chat work but the first tool call in Agent fail?

Text generation and tool calls are different capabilities. A model can produce fluent prose and still fail to emit structured tool calls, and reasoning-class models may additionally require prior thinking content to be echoed back. Test tool behaviour separately instead of inferring it from a successful chat reply.

Does a local proxy on 127.0.0.1 work as a base URL?

No. Requests originate from Cursor's servers, not your machine, and private addresses are refused. The endpoint has to be publicly reachable over HTTPS with a certificate from a trusted authority.

Does enabling HTTP/1.1 compatibility mode help?

Sometimes, and specifically on networks where a proxy terminates HTTP/2 streams. If Agent hangs mid-response at the office but works on a home connection, it is the first setting worth changing.

Related

Get API access