Cursor Agent Errors on a Custom Endpoint: How to Localise Them
Chat answers. Agent does not. That asymmetry is the most reported shape of custom-endpoint trouble in Cursor, and it is also the one most often misdiagnosed as a bad URL — which sends people off rewriting a base URL that was correct all along.
The reason the two disagree is that they are different workloads wearing the same UI. A chat turn is one request: some text in, some text out. An Agent turn is that request plus structured tool calls, plus a Server-Sent Events stream the client holds open, plus a messages array that grows with every tool result until something hits a limit. Anything in that chain can fail while text generation keeps working perfectly.
https://aicomp.ai/v1).
Create one free →
So this page is organised as triage, not as setup. Every section starts from something you can see in Cursor and ends at a cause you can confirm from the server side. Read it with your endpoint's request log open — half of these are unanswerable from the client, which is the point worth internalising before you start.
The triage table
Start here. Map what you saw to the cheapest test that distinguishes it.
| What you see | Most likely cause | The test that separates them |
|---|---|---|
| Chat works, first tool call stalls or errors | Model lacks function-calling support | Send one request with a tools payload to the same model with curl |
| Second tool call fails, first succeeded | Reasoning content not echoed back in history | Inspect the messages array your endpoint received on turn two |
| Response hangs, no error shown | Stream interrupted, or buffered instead of streamed | Watch for partial SSE frames; try HTTP/1.1 compatibility mode |
| Works for a few turns, then fails every time | Context exhausted | Start a new chat; if it works again immediately, that was it |
| Every model fails at once | MCP server conflict, or a wider outage | Disable all MCP servers, restart, retry one model |
model not found in Agent only | Different model selected than the one you added | Check the picker, not the settings page |
| Fails at the office, works at home | Proxy breaking HTTP/2 | Try another network before changing any setting |
Two of those deserve more than a table row.
Reasoning models and the second-turn failure
This one has a specific fingerprint and it wastes days when you do not know it.
Certain reasoning-class models emit thinking content alongside their reply, and they validate that this content comes back in the conversation history. The OpenAI Chat Completions format has no field for it, so a client building the next request simply drops it. The model then rejects the request: an error naming a missing reasoning field in an assistant message that carries tool calls.
Note exactly when it fires. Not on the first tool call — that one succeeds. It fails on the next turn, the one assembled from the previous turn's tool result. And it never appears in plain chat, because plain chat never produces an assistant message containing tool calls. A user testing with a sentence and declaring the endpoint healthy will never see it, which is why reports of this cluster so hard right after "it works for me".
To confirm it rather than guess at it, read the request your endpoint actually received on the failing turn and look for an assistant message with a tool_calls array. If the thinking field is absent from that message, you have your answer, and the fix is on the client side — not something a different base URL will solve.
The practical consequence for model selection is blunt: some reasoning models are fine for chat through Cursor and unreliable for Agent work. That is a property of the model and the client together, and it does not depend on which gateway sits in between. If you need Agent behaviour specifically, validate tool calls before committing — the tiers of model capability are not visible from a price table.
Context exhaustion disguised as a provider error
A long session accumulates tokens. Every tool result becomes part of the next request, and by turn ten or fifteen the messages array can be larger than the file you were editing.
When that hits the model's limit, the backend returns a resource error and Cursor surfaces something that reads like a provider failure. No configuration is wrong. Nothing about your endpoint changed. Starting a new chat fixes it immediately, which is the diagnostic: a fix that resets history and nothing else means history was the problem.
There is a related and cheaper trap worth ruling out early. The Auto model selector routes dynamically and can land on a rate-limited backend, producing failures that look like authentication trouble. Switching to a named model takes ten seconds and removes that whole class of ambiguity, so do it before you open the base URL settings at all.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-sonnet-5 | $1.00 / $5.00 | $0.48 | $9.60 |
| claude-opus-5 | $2.5 / $12.5 | $1.20 | $24.00 |
| gpt-5.6-sol | $2.5 / $15.00 | $1.35 | $27.00 |
| deepseek-v4-flash | $0.22 / $0.66 | $0.08 | $1.58 |
| gemini-3.7-flash | $0.375 / $1.875 | $0.18 | $3.60 |
One day = 180k input + 60k output on this page's workload. 20 days = 3.6M input and 1.2M output. Rates checked 2026-09-26 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
That input-dominant shape explains why this page's workload — 180k input and 60k output tokens a day, 3.6M input and 1.2M output over 20 working days — costs what it does, and why context exhaustion is an economic limit before it is a technical one. The cheapest intervention for many of these failures is not a larger model but a shorter conversation: start a new chat when the task changes, rather than carrying six tasks of history forward. The workload-based model chooser is built around exactly that trade.
When everything fails at once
If every model returns errors simultaneously, stop looking at models.
The most commonly reported cause is a single misbehaving MCP server. MCP runs at the layer shared by all model requests, so one broken server blocks everything downstream of it, including models that were working an hour ago. Cursor exposes them under Settings → MCP. Disable all of them, restart, and try one request. If the error clears, re-enable one at a time; do not re-enable them in bulk, because you will only learn that the problem is back.
Check status pages before that, though, because it is faster: an upstream incident lasts hours and is not yours to fix. And note that strict per-model tests are not the only thing worth separating here — a 401 and a stream failure can present identically in the UI while having nothing in common.
Why your endpoint's log answers questions Cursor cannot
Several useful distinctions exist only on the server. Did a request arrive at all? With which path, and did it carry a doubled version prefix? Was the model ID exactly what you typed, or something the client substituted? Did the response complete, or did the connection drop partway through a stream?
Cursor will tell you none of those. This is also why the local failures in the general endpoint checklist are worth running first — they are cheap, they do not involve Cursor, and they clear the ground.
One boundary that catches people here: requests do not originate from your machine. They come from Cursor's servers, so localhost cannot work and private addresses are refused outright. Success in a local terminal establishes connectivity from the terminal and nothing more — see why localhost fails as a base URL for the governance consequences that follow from that same fact.
FAQ
FAQ
Why does chat work but Agent fails on the first tool call?
Text generation and tool calling are separate model capabilities. A model can produce excellent prose and still fail to emit structured tool calls. Confirm it by sending a single request with a tools payload to the same model from curl — if that fails too, no Cursor setting will fix it.
Why does it fail on the second tool call rather than the first?
Reasoning-class models expect their thinking content echoed back in the conversation history. Clients using the OpenAI message format drop it, so the turn following the first tool result gets rejected. Read the messages array your endpoint received to confirm.
Why does it start failing ten minutes into a session?
Context exhaustion. Every tool result enlarges the next request until the model's window is full, and the backend error surfaces as something that looks like a provider fault. Starting a new chat resolving it immediately is the confirmation.
Can an MCP server really break every model at once?
Yes, and it is the most reported cause of simultaneous failures across all models. MCP shares the request layer, so one misbehaving server blocks everything. Disable all of them, restart, and re-enable one at a time.
Does switching off Auto actually help?
Often. Auto routes dynamically and can select a rate-limited backend, generating failures that resemble configuration errors. Naming a model removes that variable, costs nothing, and should precede any other diagnostic.
Does a smarter fix exist than changing models?
Usually it is a shorter conversation. Because the workload is input-dominant, history is both the technical limit and the cost driver. Starting a new chat per task avoids the majority of context failures.
Related
- Cursor Agent Mode with a custom endpoint — which surfaces route to your key and which never do
- Cursor OpenAI Compatible Base URL — the four URL formats that fail, and adding model IDs correctly
- Cursor API key errors — 401s, 403s and scoping, read off the status code
- Base URL not working — the seven causes, ordered by how often they actually occur