context_length_exceeded on a custom endpoint
Your client does not know how big the window is — the server does. Your client sends whatever you handed it, and the server either rejects the request or, worse, quietly trims it to fit. Both outcomes produce a valid-looking response; only one tells you what happened.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
What the error actually says
The response body carries a code, a message naming the model's limit, and often the count you sent:
{
"error": {
"message": "This model's maximum context length is 1048576 tokens. "
"You requested 1203433 tokens in the messages.",
"type": "invalid_request_error",
"param": "messages",
"code": "context_length_exceeded"
}
}
Two numbers, and the gap between them is your whole diagnosis. Everything else — how big you thought the file was, how many messages you remember sending — is unreliable, because the count includes the system prompt, the entire history, every tool result appended during the session, and every retrieved document. None of that is visible from where you are sitting.
Two independent limits
The input window and the maximum output are separate caps, and the failure modes differ:
- Window exceeded — the request is rejected before inference. You get an error and, usefully, the number.
- Output cap reached — the request runs and stops early. You get a valid response with
finish_reasonoflength, truncated mid-thought, and no error at all.
Among the models in our catalogue where both figures are published, the input window is between 3.1× and 8.2× the maximum single output. So on generation-heavy work the output cap is the binding constraint even though nobody talks about it.
| Model | Input window | Max output | Window ÷ output | Rate in / out per 1M | One full-window request |
|---|---|---|---|---|---|
gpt-5.6-luna | 1,050,000 | 128,000 | 8.2× | $0.1 / $0.6 | $0.1818 |
claude-haiku-4-5-20251001 | 200,000 | 64,000 | 3.1× | $0.5 / $2.5 | $0.26 |
claude-sonnet-5 | 1,000,000 | 128,000 | 7.8× | $1 / $5 | $1.64 |
gpt-5.6-terra | 1,050,000 | 128,000 | 8.2× | $1 / $6 | $1.818 |
claude-opus-5 | 1,000,000 | 128,000 | 7.8× | $2.5 / $12.5 | $4.1 |
gpt-5.6-sol | 1,050,000 | 128,000 | 8.2× | $2.5 / $15 | $4.545 |
claude-fable-5 | 1,000,000 | 128,000 | 7.8× | $5 / $25 | $8.2 |
Only 7 of the models we track publish both figures — the rest leave one or both unspecified, which is itself worth knowing before you plan around them. Rates checked 2026-09-26; gateway rates move with upstream promotions.
Stop estimating, start measuring
The only authoritative count of what you are sending is the one the server reported on your last successful call. Make it non-streamed, because the usage field is always populated there, and read prompt_tokens.
# The only authoritative count of what you are sending is the one the
# server reported on your last successful call. Estimates are always wrong
# because system prompts, tool results and retrieved documents are all
# counted and none of them are visible in your editor.
last = client.chat.completions.create(
model="claude-sonnet-5",
messages=messages,
stream=False, # non-streamed: the usage field is always populated
)
used = last.usage.prompt_tokens
print(f"last successful call: {used} input tokens")
print(f"implied headroom : {WINDOW - used}") # WINDOW = the model's published value
# Then, when it finally fails, read the number the server says rather than
# arguing with it -- the error body carries the model's own limit.
# If the response instead comes back short and complete, nothing was rejected:
# something upstream truncated, and finish_reason is the tell.
print("finish_reason:", last.choices[0].finish_reason) # "length" = capped by max_tokens
That gives you the number to reason with, and it gives you a baseline to watch. Plot it across a session and the shape tells you which part of the context is growing — a steady linear climb is history, a jump is a tool result or a retrieved document.
The version that never errors
Worth repeating, because it is the one that costs the most: clients infer the window from the model name. Serve a model under a custom id and the client may assume a small window and trim your prompt before it is sent. No rejection, no error message, a perfectly valid response — built on less context than you provided.
If your client has a model configuration section, fill in the window and the maximum output explicitly for every custom id. If it does not, treat any unexplained quality drop on long inputs as this failure until proven otherwise. Our /v1/models page covers where the id comes from in the first place.
The order to fix things in
- Compact the conversation. Replace history with a summary of decisions made. Usually the single largest recovery.
- Trim tool results. A file read or a shell output is often tens of thousands of tokens for a few hundred that mattered. Keep the relevant extract.
- Drop retrieved documents that did not change the answer. Retrieval returns candidates, not all of which are needed in context.
- Set
max_tokensdeliberately. An unnecessarily large value does not cause the error but it does cap output incorrectly the other way. - Then change models. Bigger window, same task — and re-measure, because a larger window on a per-token basis also means a larger worst-case request.
Steps one to three also cut cost, which is the reason to do them before the model swap rather than after. Across the 216 models in our cleaned catalogue the median output-to-input ratio is 4.0×, and on agent-shaped work the input side dominates the invoice — the input/output breakdown shows how to tell which half yours sits on.
Related guides
- One endpoint for every tool — the hub for tool-specific setup
- Reading /v1/models — where the id and the window hint come from
- Streaming and usage — getting the token count you need to see this coming
- Output vs input pricing — which side of the sum your workload lives on
- Cheapest model by workload — repricing the same task
- Claude Code 429 — the other limit that stops a session mid-flight
Context-limit failures and what each one is actually telling you
| Symptom | What is actually happening | How to confirm | Fix |
|---|---|---|---|
| context_length_exceeded on a small task | History, tool results and retrieved documents are all counted and re-sent every turn. | Read prompt_tokens from one non-streamed call. | Compact history and trim tool results before retrying. |
| Responses arrive truncated with no error | finish_reason is length: the output cap stopped generation, not the window. | Print finish_reason on the failing response. | Raise max_tokens within the model's published output limit. |
| Long files are silently shortened | The client inferred a small window from an unrecognised model id and trimmed the prompt. | Compare the tokens you sent against what the server counted. | Set the window explicitly in the client's model configuration. |
| The error appears only after many turns | Context grows monotonically across a session; nothing shrinks it automatically. | Plot prompt_tokens per turn. | Compact on a schedule, not when it breaks. |
| A bigger-window model did not help | The binding constraint was the output cap, or the context was never the problem. | Check finish_reason and the sent count in the error body. | Fix the constraint the evidence names. |
| Cost per session climbs faster than usage | Longer context is re-billed on every turn, so the same task gets more expensive as the session grows. | Track cost per turn alongside prompt_tokens. | Compact earlier; the saving compounds over the session. |
When this is the wrong move
Four moves that fix the wrong limit
| Situation | Why it breaks | Do this instead |
|---|---|---|
| You are about to switch models before measuring | Without the sent count you cannot tell whether the window was the binding constraint at all. | Read the error body and one non-streamed usage object first. |
| You are about to raise max_tokens to fix a rejection | A rejection is the input window; max_tokens governs output. They are independent. | Fix the constraint the error names. |
| You are about to rely on the client's inferred window | An unrecognised id gets a conservative default and prompts are trimmed silently. | State the window explicitly for every custom id. |
| You are about to treat a truncated response as a complete one | finish_reason of length means the model was stopped, not that it finished. | Assert on finish_reason before using the output. |
Rolling it out without finding out the hard way
The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:
- Prove it on one long session with prompt_tokens and finish_reason logged per turn. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
- Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
- Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
- Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.
Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.
FAQ
Why did I get this error when my files are small?
Because the input count is not your files. It is the system prompt, the full conversation history, every tool result appended during the session, and every retrieved document — all re-sent on every turn. A session that opened with a modest prompt can be carrying hundreds of thousands of tokens by turn thirty without anything visibly growing. The error body usually names both the model's limit and the count you sent; compare those two numbers rather than reasoning about file sizes.
Is the input window the same as the maximum output?
No, they are independent limits, and conflating them sends people to the wrong fix. Among the models in our catalogue where both are published, the input window is between 3.1× and 8.2× the maximum single output. A generation task can therefore sit comfortably inside the window and still be cut off by the output cap. The tell is finish_reason: a value of length means the output cap stopped the model, not the window.
How do I find out how much context I am actually sending?
Ask the server, not your editor. Make one non-streamed call with the same messages and read usage.prompt_tokens — non-streamed because the usage field is always populated there. Streamed calls need stream_options and can lose the count entirely, which is exactly why our streaming page treats a missing usage chunk as unknown rather than zero.
Why does my client truncate long files without any error?
Because it inferred the window from the model name. Serve a model under a custom id and the client may fall back to a conservative default and trim your prompt before sending it — no rejection, no error, a valid response built on less context than you thought. Where the client exposes a model configuration section, enter the window and maximum output explicitly for every custom id.
Does a bigger window cost more?
Only if you fill it. Tokens are billed per token, so an unfilled large window is free. The table below prices the ceiling rather than the typical case: a single request that uses the whole window and the whole output cap. Those numbers are useful as a worst case per request, not as a forecast.
Should I switch to a model with a bigger window?
Last, not first. Compacting the conversation, trimming tool results and dropping retrieved documents that did not change the answer usually recovers more headroom than a model swap, and it also cuts cost. Switch when the task genuinely needs the context and you have already removed what you can. Our workload-based comparison prices the same task across models.