OpenAI-compatible error codes: 401 vs 404

Most wasted debugging in this area comes from one confusion: treating a 401 and a 404 as the same kind of problem. Authentication happens before routing, so a 401 says your credential was rejected and nothing about your path, while a 404 says the credential passed and the routing is wrong. Getting it backwards sends you to edit the one thing that was correct.

In short: Authentication happens before routing: a 401 means the key was rejected and says nothing about the path, while a 404 means the key passed. Three unrelated conditions all return 429.
You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment
Confirm the key works first. One command, no SDK, costs nothing:
curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.

The matrix

StatusWhat it actually meansCommon causeHow to confirm in one stepWhat to do
401Credential rejected before routingWrong key, stale environment variable, or the key passed to the wrong parameter.Send the same request with no key. Still 401 → the path is fine and the key is not.Print the key prefix the client holds. Never rotate a key that was never read.
403Authenticated but not permittedThe key is valid and the request reached the service — this one is about entitlement, not identity.Compare against a model you know works. Same key, one model fails → entitlement.Check what that key is scoped to. Do not treat this as an auth bug.
404Routed wrong, credential finePath shape off — a doubled or missing /v1 — or the endpoint does not implement that route.curl /v1/models. A list comes back → key and path root are correct.Match the path shape that works in curl. Do not touch the key.
422Request rejected on its shapeA field the endpoint does not accept, or a parameter combination it will not serve.Remove parameters one at a time from the smallest failing request.Send the minimum viable body, then add fields back.
429Rate limited — three different thingsPer-minute ceiling, period allowance, or a spend cap. Reset behaviour differs by orders of magnitude.Look for a reset time in the message or a retry-after header.Classify before acting. Only per-minute limits are worth waiting out.
500 / 502 / 503Upstream failure, not your requestThe failure is on the other side of the connection. Your request may or may not have been billed.Retry once against /v1/models. If that works, the route is fine.Back off with jitter. Do not retry in a tight loop.
timeoutNo answer within your budgetYour timeout is shorter than the request needs, or the connection is being reset.Compare your timeout against the latency of a request that succeeds.Set timeout and retry count together — they compound.
200 with no contentSuccess status, empty resultCapability negotiation failed silently, or the response body is not the shape you parsed.Print the raw body and count the choices array.Assert on the parsed shape, not on the status code.

Status meanings are the HTTP ones. The causes and confirmation steps are specific to routing an OpenAI-compatible client at an endpoint that is not the one it defaults to.

One script, every class

Run this before you start guessing. It establishes a baseline, then separates each class in one request.

#!/usr/bin/env bash
# One script, every error class. Run it before you start guessing.
BASE="https://aicomp.ai/v1"
KEY="$GATEWAY_API_KEY"
MODEL="$1"   # pass a model ID to test the last case

code () { curl -sS -o /tmp/body.json -w "%{http_code}" "$1" -H "Authorization: Bearer $KEY" ${2:+-d "$2" -H 'Content-Type: application/json'}; }

echo "no key      -> $(curl -sS -o /dev/null -w '%{http_code}' "$BASE/models")   (expect 401)"
echo "bad key     -> $(curl -sS -o /dev/null -w '%{http_code}' "$BASE/models" -H 'Authorization: Bearer nope')   (expect 401)"
echo "good key    -> $(code "$BASE/models")   (expect 200 -> key and path are fine)"
echo "bad path    -> $(code "$BASE/v1/models")   (404 here = path shape is wrong)"
echo "bad model   -> $(code "$BASE/chat/completions" '{"model":"does-not-exist","messages":[{"role":"user","content":"hi"}]}')"

# The two that cost the most money, because both look like success:
echo
echo "--- 200 but empty: read the body, not just the status"
curl -sS "$BASE/chat/completions" -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -d "{\"model\":\"$MODEL\",\"messages\":[{\"role\":\"user\",\"content\":\"Say: ok\"}]}" \
  | python3 -c "import json,sys; d=json.load(sys.stdin); print('choices:', len(d.get('choices') or [])); print('usage  :', d.get('usage'))"

echo
echo "--- 429: ask for the reset time, do not assume one"
curl -sS -D - -o /dev/null "$BASE/chat/completions" -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -d "{\"model\":\"$MODEL\",\"messages\":[{\"role\":\"user\",\"content\":\"Say: ok\"}]}" \
  | grep -iE "^(HTTP/|retry-after|x-ratelimit|ratelimit)" 

The first three lines do the heavy lifting. A request with no key that returns 401 tells you the path resolves and the service is reachable — which means every subsequent 401 is about the credential, and every 404 is about routing. That single baseline removes the ambiguity that makes this error class slow.

The two failures that look like success

Everything in the matrix above announces itself. These two do not, which is why they cost more than any 500.

A 200 with an empty result. A status code reports transport success, not semantic success. The common cause is capability negotiation failing silently: a model that does not support what you asked for returns prose instead of an error, and a client expecting structured output reads nothing. Nothing in your monitoring fires.

A request that was billed but raised. Anything that reached the model consumed input tokens. A timeout, a 5xx, or a cancelled stream can all have been billed while surfacing only as an exception. You cannot tell from the status code — you have to reconcile your request log against the usage record.

Cost note. These two are why retry configuration is a cost setting rather than a reliability setting. Every retry is another request that may be billed whether or not it succeeds, and every silent success is a request you paid for and did not use.

Reading a 429 properly

Three conditions return the same status and behave nothing alike:

ConditionResets afterWaiting helps?
Per-minute request or token ceilingAbout a minuteYes — this is what waiting is for
Period allowanceHours to daysOnly if the work can wait that long
Spend capDoes not reset on its ownNo — waiting is waiting for nothing

The message often names a reset time. When it does not, that absence is itself the signal — a limit with no reset time is usually not a time-based limit at all.

What it costs

Models behind these errors — gateway rate vs official list
ModelGateway rate
in / out per 1M tokens
Official list
in / out per 1M tokens
Diff
gpt-5.6-luna$0.1 / $0.6$0.2 / $1.250%
claude-sonnet-5$1 / $5$2 / $1050%
deepseek-v4-flash$0.22 / $0.66— / ——

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget. Across the 216 models in our cleaned catalogue the median output-to-input ratio is 4.0×.

Error-handling failures and what each one is actually telling you

SymptomWhat is actually happeningHow to confirmFix
401 and 404 get treated as one problemAuthentication precedes routing, so the two localise the fault to different places.Send one request with no key at all and read what comes back.Fix the credential on a 401 and the path on a 404. Never both at once.
A 429 gets waited out for an hourThree conditions share the status; only one of them is time-based at that scale.Look for a reset time in the message or a retry-after header.Classify first. A spend cap never clears by waiting.
A 200 with no content goes unnoticedStatus codes report transport success. An empty choices array is still a failure.Count the choices array on one response.Assert on the parsed shape, not on the status.
Failed requests are assumed to be freeAnything that reached the model consumed input tokens, including requests that ended in an exception.Reconcile your request log against the usage record.Treat the difference as real spend and put a ceiling on retries.
Retries are added until the errors stopRetries consume the same window that refused you, and each one may be billed.Count upstream attempts per user request.Cap the retry budget and add backoff with jitter.
Every error becomes a key rotationKey rotation fixes a 401 and nothing else, and it invalidates working deployments on the way.Compare behaviour with no key and with a known-good one.Rotate only when the baseline says the credential is the fault.

When this is the wrong move

Three habits that make error handling expensive

SituationWhy it breaksDo this instead
You are about to retry every error uniformlyAuthentication, not-found and content-policy errors are marked non-retryable: the resent request is identical, so the answer is identical.Split the policy along the transient-versus-permanent line.
You are about to rotate a key as a first responseRotation fixes one status code, and it breaks every deployment using the old key while you check.Establish the no-key baseline first.
You are treating a 200 as proof the integration worksTransport success is not semantic success, and an empty result is the most expensive failure on this page.Assert on the parsed shape.

Rolling it out without finding out the hard way

The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:

  1. Prove it on one request per error class, run against a non-critical route. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
  2. Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
  3. Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
  4. Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.

Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.

FAQ

What is the real difference between a 401 and a 404?

The difference localises the fault, which is why it is the most useful distinction on this page. Authentication happens before routing. A 401 means the credential was rejected, so the request never reached a model and the result tells you nothing about your path. A 404 means the credential passed and the routing is wrong. Getting these backwards sends people to edit a base URL that was correct, or to rotate a key that was never the problem — and both feel productive while achieving nothing.

Is a 429 always about how much I have used?

No, and this is where the most time gets wasted. Three distinct conditions return it: a per-minute ceiling, a period allowance, and a spend cap. Their reset behaviours differ by orders of magnitude — roughly a minute, hours to days, and never. Look for an explicit reset time in the message or a retry-after header before you decide whether to wait or to change something. Waiting out a spend cap is waiting for nothing.

Why did I get a 200 with nothing in it?

Because a status code reports transport success, not semantic success. The usual causes are capability negotiation failing silently — a model that does not support the feature you asked for returns prose instead of an error — or a response body that is not the shape your parser expects. Both are more expensive than a 500, because nothing in your monitoring fires. Print the raw body and count the choices array; that single check finds it.

Should I retry a 500?

Yes, but with a backoff rather than in a loop. A 5xx says the failure is on the other side of the connection, so a later identical request may succeed. What a 5xx does not tell you is whether the request that failed was billed — if it reached the model, its input tokens were consumed. That is why retry budgets need a ceiling rather than a hope.

How do I know whether a failed request cost me money?

You usually cannot tell from the status code alone. Anything that reached the model consumed input tokens, and a timeout or a 5xx may have done exactly that while surfacing only as an exception. The only reliable answer is server-side: reconcile your own request log against the usage record, and treat the difference as real spend. This is the gap that makes retry configuration a cost setting rather than a reliability setting.

What is the single cheapest guard I can add?

Assert on the parsed response shape rather than on the status code. One line checking that the choices array is non-empty, and one checking that a usage block came back, converts both silent failures on this page into immediate ones at the point they happen. Nearly every expensive debugging session in this area starts with a 200 that was not actually a success.

Get API access