OpenAI-compatible error codes: 401 vs 404
Most wasted debugging in this area comes from one confusion: treating a 401 and a 404 as the same kind of problem. Authentication happens before routing, so a 401 says your credential was rejected and nothing about your path, while a 404 says the credential passed and the routing is wrong. Getting it backwards sends you to edit the one thing that was correct.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
The matrix
| Status | What it actually means | Common cause | How to confirm in one step | What to do |
|---|---|---|---|---|
401 | Credential rejected before routing | Wrong key, stale environment variable, or the key passed to the wrong parameter. | Send the same request with no key. Still 401 → the path is fine and the key is not. | Print the key prefix the client holds. Never rotate a key that was never read. |
403 | Authenticated but not permitted | The key is valid and the request reached the service — this one is about entitlement, not identity. | Compare against a model you know works. Same key, one model fails → entitlement. | Check what that key is scoped to. Do not treat this as an auth bug. |
404 | Routed wrong, credential fine | Path shape off — a doubled or missing /v1 — or the endpoint does not implement that route. | curl /v1/models. A list comes back → key and path root are correct. | Match the path shape that works in curl. Do not touch the key. |
422 | Request rejected on its shape | A field the endpoint does not accept, or a parameter combination it will not serve. | Remove parameters one at a time from the smallest failing request. | Send the minimum viable body, then add fields back. |
429 | Rate limited — three different things | Per-minute ceiling, period allowance, or a spend cap. Reset behaviour differs by orders of magnitude. | Look for a reset time in the message or a retry-after header. | Classify before acting. Only per-minute limits are worth waiting out. |
500 / 502 / 503 | Upstream failure, not your request | The failure is on the other side of the connection. Your request may or may not have been billed. | Retry once against /v1/models. If that works, the route is fine. | Back off with jitter. Do not retry in a tight loop. |
timeout | No answer within your budget | Your timeout is shorter than the request needs, or the connection is being reset. | Compare your timeout against the latency of a request that succeeds. | Set timeout and retry count together — they compound. |
200 with no content | Success status, empty result | Capability negotiation failed silently, or the response body is not the shape you parsed. | Print the raw body and count the choices array. | Assert on the parsed shape, not on the status code. |
Status meanings are the HTTP ones. The causes and confirmation steps are specific to routing an OpenAI-compatible client at an endpoint that is not the one it defaults to.
One script, every class
Run this before you start guessing. It establishes a baseline, then separates each class in one request.
#!/usr/bin/env bash
# One script, every error class. Run it before you start guessing.
BASE="https://aicomp.ai/v1"
KEY="$GATEWAY_API_KEY"
MODEL="$1" # pass a model ID to test the last case
code () { curl -sS -o /tmp/body.json -w "%{http_code}" "$1" -H "Authorization: Bearer $KEY" ${2:+-d "$2" -H 'Content-Type: application/json'}; }
echo "no key -> $(curl -sS -o /dev/null -w '%{http_code}' "$BASE/models") (expect 401)"
echo "bad key -> $(curl -sS -o /dev/null -w '%{http_code}' "$BASE/models" -H 'Authorization: Bearer nope') (expect 401)"
echo "good key -> $(code "$BASE/models") (expect 200 -> key and path are fine)"
echo "bad path -> $(code "$BASE/v1/models") (404 here = path shape is wrong)"
echo "bad model -> $(code "$BASE/chat/completions" '{"model":"does-not-exist","messages":[{"role":"user","content":"hi"}]}')"
# The two that cost the most money, because both look like success:
echo
echo "--- 200 but empty: read the body, not just the status"
curl -sS "$BASE/chat/completions" -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
-d "{\"model\":\"$MODEL\",\"messages\":[{\"role\":\"user\",\"content\":\"Say: ok\"}]}" \
| python3 -c "import json,sys; d=json.load(sys.stdin); print('choices:', len(d.get('choices') or [])); print('usage :', d.get('usage'))"
echo
echo "--- 429: ask for the reset time, do not assume one"
curl -sS -D - -o /dev/null "$BASE/chat/completions" -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
-d "{\"model\":\"$MODEL\",\"messages\":[{\"role\":\"user\",\"content\":\"Say: ok\"}]}" \
| grep -iE "^(HTTP/|retry-after|x-ratelimit|ratelimit)"
The first three lines do the heavy lifting. A request with no key that returns 401 tells you the path resolves and the service is reachable — which means every subsequent 401 is about the credential, and every 404 is about routing. That single baseline removes the ambiguity that makes this error class slow.
The two failures that look like success
Everything in the matrix above announces itself. These two do not, which is why they cost more than any 500.
A 200 with an empty result. A status code reports transport success, not semantic success. The common cause is capability negotiation failing silently: a model that does not support what you asked for returns prose instead of an error, and a client expecting structured output reads nothing. Nothing in your monitoring fires.
A request that was billed but raised. Anything that reached the model consumed input tokens. A timeout, a 5xx, or a cancelled stream can all have been billed while surfacing only as an exception. You cannot tell from the status code — you have to reconcile your request log against the usage record.
Reading a 429 properly
Three conditions return the same status and behave nothing alike:
| Condition | Resets after | Waiting helps? |
|---|---|---|
| Per-minute request or token ceiling | About a minute | Yes — this is what waiting is for |
| Period allowance | Hours to days | Only if the work can wait that long |
| Spend cap | Does not reset on its own | No — waiting is waiting for nothing |
The message often names a reset time. When it does not, that absence is itself the signal — a limit with no reset time is usually not a time-based limit at all.
What it costs
| Model | Gateway rate in / out per 1M tokens | Official list in / out per 1M tokens | Diff |
|---|---|---|---|
| gpt-5.6-luna | $0.1 / $0.6 | $0.2 / $1.2 | 50% |
| claude-sonnet-5 | $1 / $5 | $2 / $10 | 50% |
| deepseek-v4-flash | $0.22 / $0.66 | — / — | — |
Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.
Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget. Across the 216 models in our cleaned catalogue the median output-to-input ratio is 4.0×.
Error-handling failures and what each one is actually telling you
| Symptom | What is actually happening | How to confirm | Fix |
|---|---|---|---|
| 401 and 404 get treated as one problem | Authentication precedes routing, so the two localise the fault to different places. | Send one request with no key at all and read what comes back. | Fix the credential on a 401 and the path on a 404. Never both at once. |
| A 429 gets waited out for an hour | Three conditions share the status; only one of them is time-based at that scale. | Look for a reset time in the message or a retry-after header. | Classify first. A spend cap never clears by waiting. |
| A 200 with no content goes unnoticed | Status codes report transport success. An empty choices array is still a failure. | Count the choices array on one response. | Assert on the parsed shape, not on the status. |
| Failed requests are assumed to be free | Anything that reached the model consumed input tokens, including requests that ended in an exception. | Reconcile your request log against the usage record. | Treat the difference as real spend and put a ceiling on retries. |
| Retries are added until the errors stop | Retries consume the same window that refused you, and each one may be billed. | Count upstream attempts per user request. | Cap the retry budget and add backoff with jitter. |
| Every error becomes a key rotation | Key rotation fixes a 401 and nothing else, and it invalidates working deployments on the way. | Compare behaviour with no key and with a known-good one. | Rotate only when the baseline says the credential is the fault. |
When this is the wrong move
Three habits that make error handling expensive
| Situation | Why it breaks | Do this instead |
|---|---|---|
| You are about to retry every error uniformly | Authentication, not-found and content-policy errors are marked non-retryable: the resent request is identical, so the answer is identical. | Split the policy along the transient-versus-permanent line. |
| You are about to rotate a key as a first response | Rotation fixes one status code, and it breaks every deployment using the old key while you check. | Establish the no-key baseline first. |
| You are treating a 200 as proof the integration works | Transport success is not semantic success, and an empty result is the most expensive failure on this page. | Assert on the parsed shape. |
Rolling it out without finding out the hard way
The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:
- Prove it on one request per error class, run against a non-critical route. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
- Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
- Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
- Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.
Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.
FAQ
What is the real difference between a 401 and a 404?
The difference localises the fault, which is why it is the most useful distinction on this page. Authentication happens before routing. A 401 means the credential was rejected, so the request never reached a model and the result tells you nothing about your path. A 404 means the credential passed and the routing is wrong. Getting these backwards sends people to edit a base URL that was correct, or to rotate a key that was never the problem — and both feel productive while achieving nothing.
Is a 429 always about how much I have used?
No, and this is where the most time gets wasted. Three distinct conditions return it: a per-minute ceiling, a period allowance, and a spend cap. Their reset behaviours differ by orders of magnitude — roughly a minute, hours to days, and never. Look for an explicit reset time in the message or a retry-after header before you decide whether to wait or to change something. Waiting out a spend cap is waiting for nothing.
Why did I get a 200 with nothing in it?
Because a status code reports transport success, not semantic success. The usual causes are capability negotiation failing silently — a model that does not support the feature you asked for returns prose instead of an error — or a response body that is not the shape your parser expects. Both are more expensive than a 500, because nothing in your monitoring fires. Print the raw body and count the choices array; that single check finds it.
Should I retry a 500?
Yes, but with a backoff rather than in a loop. A 5xx says the failure is on the other side of the connection, so a later identical request may succeed. What a 5xx does not tell you is whether the request that failed was billed — if it reached the model, its input tokens were consumed. That is why retry budgets need a ceiling rather than a hope.
How do I know whether a failed request cost me money?
You usually cannot tell from the status code alone. Anything that reached the model consumed input tokens, and a timeout or a 5xx may have done exactly that while surfacing only as an exception. The only reliable answer is server-side: reconcile your own request log against the usage record, and treat the difference as real spend. This is the gap that makes retry configuration a cost setting rather than a reliability setting.
What is the single cheapest guard I can add?
Assert on the parsed response shape rather than on the status code. One line checking that the choices array is non-empty, and one checking that a usage block came back, converts both silent failures on this page into immediate ones at the point they happen. Nearly every expensive debugging session in this area starts with a 200 that was not actually a success.