Claude Code 429: which limit you actually hit

"Rate limit reached" is a wrapper message, and more than one underlying condition produces it. The differences matter because the reset behaviours are three orders of magnitude apart: per-minute limits clear in about a minute, subscription periods take hours, and a spend cap never clears on its own. Guessing which one you hit is what turns a five-minute problem into an afternoon.

In short: 'Rate limit reached' is a wrapper message covering per-minute limits, period allowances and spend caps, whose reset times differ by orders of magnitude. A minimal request whose token counters are all zero was rejected before inference, not throttled by volume.
You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment
Confirm the key works first. One command, no SDK, costs nothing:
curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.

Why the message misleads

The CLI reports one sentence for several distinct conditions. That is reasonable as user interface design and unhelpful as a diagnostic, because the conditions it collapses have almost nothing in common. A request rejected because it was classified as long-context, and a request rejected because your account hit a per-minute ceiling, look identical in the terminal and call for completely different responses.

So the first move is always the same: stop debugging inside the session that is failing. A long session is itself one of the variables you are trying to measure. Run the smallest request that can fail.

# 1. Smallest possible reproduction. Do not debug inside a long session --
#    the session itself is one of the variables you are trying to measure.
claude -p --output-format json "Reply with exactly OK."

# Read three fields out of the JSON:
#   input_tokens     = 0  -> the request never reached inference
#   output_tokens    = 0  -> same
#   duration_api_ms  = 0  -> rejected at the entry point, not throttled mid-flight
#
# All three at zero points at authentication, routing or a pre-flight check --
# NOT at "you ran out of quota this week".

That single check — are the token counters zero — separates "rejected at the door" from "throttled by volume" in about ten seconds, and it prevents most of the wasted effort that follows this error.

Capture what the API actually said

The wrapper is generic on purpose. The underlying error is specific, and it is sitting in the debug output.

# 2. Capture what the CLI is actually being told. The wrapper message is
#    generic on purpose; the underlying error is not.
claude -p \
  --debug api \
  --debug-file /tmp/claude-debug.txt \
  --output-format json \
  "Reply with exactly OK."

# Then search the capture for the real error rather than reading the CLI output:
rg -n "rate_limit_error|Extra usage|long context|overloaded" /tmp/claude-debug.txt

# What you are looking for is the difference between:
#   429 rate_limit_error: "Extra usage is required for long context requests."
#       -> this request was classified as long-context. It is about the request.
#   429 "This request would exceed your account's rate limit"
#       -> this is about your account's per-minute or per-day allowance.

The distinction you are looking for is between an error about this request and an error about your account. "Extra usage is required for long context requests" is the first kind — the request was classified as long-context. "This request would exceed your account's rate limit" is the second. Reading them as the same thing is the most common way to spend an hour on the wrong fix.

One observed case is worth knowing because it is genuinely counter-intuitive: a model string carrying an extended-context marker can cause the request to be treated as long-context, and the CLI surfaces the resulting rejection as the generic rate-limit message. If your configuration names a model with an extended-context suffix and you are seeing this error with usage left over, that is the first thing to check. We flag this as observed community behaviour rather than documented product behaviour — verify the model string your own configuration is sending before acting on it.

Classify by reset behaviour, not by message

Once you have the underlying error, the useful question is how long this lasts. That answer alone tells you whether to wait or to change something.

Limit typeResets afterThe right response
Requests per minuteAbout a minuteWait, then reduce concurrency. Do not restructure anything.
Tokens per minuteAbout a minuteWait, then shrink the request. This is the one agent loops hit.
Subscription period allowanceHours to daysWait, or move that work to usage-based access.
Spend capDoes not reset by itselfRaise the cap. Waiting is waiting for nothing.

Reset windows are characteristic of each limit type rather than fixed values, and they vary by access path and plan. Use the classification to pick a response, and confirm the actual reset time from the message or your console rather than from this table.

Why agent loops hit token limits specifically

An agent loop resends context on every turn. That makes the cost of a session roughly turns × context, not turns — which is why a session that was fine at turn three starts failing at turn twenty with no change in what you asked for.

Context per turnTurnsInput total
context × turns
Output totalCost of the session
claude-sonnet-5
20k10200k12k$0.26
60k201,200k24k$1.32
120k303,600k36k$3.78
200k408,000k48k$8.24

Computed from our catalogue rate for claude-sonnet-5 — gateway rate, checked 2026-09-26. The input column is context × turns, which is the shape every agent loop produces. Output assumes ~1.2k tokens per turn.

Cost note. The multiplier is the whole story. Halving context per turn halves the session cost; halving the number of turns does the same. But shrinking context is usually the cheaper lever, because turns are what you are doing and context is what you are carrying.

Shrink the request rather than the session

# 3. Shrink the request, not the session. Agent loops resend context on
#    every turn, so the cost is roughly turns x context, not turns.

# Keep build output and dependencies out of the indexed context entirely:
# .claudeignore
node_modules/
dist/
build/
*.lock
*.min.js
coverage/

# Point subagents at a lighter model -- the work they do rarely needs the
# most expensive one. (Environment variable; confirm the name against your
# own CLI version before relying on it.)
export CLAUDE_CODE_SUBAGENT_MODEL=<your-lightweight-model-id>

# Then verify the change took effect with a fresh minimal request:
claude -p --output-format json "Reply with exactly OK." 

Two of those attack the multiplier. Excluding build output and dependency directories from the indexed context removes tokens that were never informative to begin with. Pointing subagents at a lighter model removes the most expensive per-turn calls without changing what the main loop does — the work subagents do rarely needs the largest model.

Verify both the same way: re-run the minimal request and check the token fields. A change you have not measured is a change you are assuming.

What it costs

Claude models used with Claude Code — gateway rate vs official list
ModelGateway rate
in / out per 1M tokens
Official list
in / out per 1M tokens
Diff
claude-sonnet-5$1 / $5$2 / $1050%
claude-haiku-4-5$0.5 / $2.5$1 / $550%
claude-opus-5$2.5 / $12.5$5 / $2550%

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget. Model choice matters more here than anywhere else on this site, because the loop resends context on every turn.

429s and what each one is actually telling you

SymptomWhat is actually happeningHow to confirmFix
Quota looks fine but every request failsThe message is a wrapper; the request may be rejected on its own characteristics rather than on your allowance.Run a minimal request with JSON output and check whether the token counters are zero.Capture the underlying error with debug output and act on that, not on the wrapper.
It started failing at turn twenty, not turn threeAgent loops resend context every turn, so cost is turns × context and grows within a session.Compare context size at the turn it started failing against the turn it worked.Shrink context — exclude build output, and split the task.
Waiting an hour changed nothingA spend cap does not reset on its own. You waited for a limit that was never time-based.Check whether the message names a reset time at all.Raise the cap, or move the work to a different access path.
It fails on one model and not anotherLong-context classification and per-model rate limits both vary by model, so this is a signal about the request.Run the identical minimal request against both models.Pick the model whose context and rate envelope fits the request.
Subagents are the expensive partEvery subagent call is a full billed request, and they inherit the main loop's context shape.Count calls by model if your tooling exposes it.Point subagents at a lighter model.
Restarting the CLI fixed it once and never againRestarting clears session context, so it works when the cause was accumulated context and nothing else.Check whether the fix survives a second long session.Fix the context, not the process.

When this is the wrong move

Three responses that make a 429 worse

SituationWhy it breaksDo this instead
You are about to retry in a tight loopRetries consume the same window that just refused you, and each one is a billed request that may also be resending full context.Back off, then shrink the request rather than resending it.
You are about to upgrade because of one errorAn upgrade fixes account-level ceilings. It does not fix a request that was rejected on its own characteristics.Classify first — the diagnostic order above takes about two minutes.
You plan to raise concurrency to get through fasterPer-minute limits are the constraint; more concurrency consumes the same window faster.Reduce concurrency and add it back while watching the error rate.

Rolling it out without finding out the hard way

The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:

  1. Prove it on one minimal request with JSON output, before you change any configuration. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
  2. Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
  3. Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
  4. Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.

Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.

FAQ

Why do I get <em>Rate limit reached</em> when my dashboard still shows quota left?

Because the message the CLI shows is a wrapper, and more than one underlying condition produces it. A request can be rejected on its own characteristics — most commonly being classified as a long-context request — while your account still has plenty of allowance. The wrapper collapses those into one sentence, which is why the dashboard and the terminal appear to disagree. The way out is to stop reading the wrapper: run one minimal request with JSON output and look at the actual fields, then capture the underlying API error with debug output to a file.

How can I tell whether the request was rejected before it reached the model?

Run a minimal prompt with JSON output and read three fields: input tokens, output tokens, and API duration. If all three are zero, the request did not reach inference — it was stopped at authentication, routing, or a pre-flight check. That is a categorically different problem from running out of allowance, and the fixes do not overlap. This one check prevents most of the wasted effort in this area.

What does <em>Extra usage is required for long context requests</em> mean?

It means this particular request was classified as long-context and the current access path does not cover it. The classification is about the request, not the account — which is why it can appear when overall usage looks low. There are documented cases of this being triggered by a model string carrying an extended-context marker, with the CLI surfacing it only as the generic rate-limit message. Treat it as a signal to look at how the request is being assembled, not as evidence that your plan has been cut.

Should I wait, or change something?

It depends entirely on which limit you hit, and the reset behaviours are three orders of magnitude apart. Per-minute request and token limits reset within roughly a minute, so waiting is the right answer. Subscription-period limits reset over hours. A spend cap does not reset on its own at all — waiting for it is waiting for nothing. Classifying before acting is the whole value of the diagnostic order on this page; guessing costs hours.

Why does my context keep growing within one session?

An agent loop resends context on every turn, so the cost of a session is roughly turns multiplied by context, not turns. Twenty turns at 60k context is 1.2M input tokens even if your last prompt was one sentence. This is why a session that worked fine at turn three starts failing at turn twenty with no change in what you are asking. Excluding build output from the indexed context and pointing subagents at a lighter model both attack the multiplier rather than the symptom.

Does switching to a usage-based endpoint remove these limits?

It changes which limits apply. Subscription access carries period-based allowances that reset on a schedule; usage-based access is bounded by what you have paid for and by per-minute rate limits. You trade a quota that resets on someone else's schedule for one that resets when you add credit — but the per-minute limits and the long-context classification can still apply, so this is a change in which constraints bind, not a removal of constraints.

Get API access