Base URL not working: seven causes, in order

A base URL that is never read and a base URL that is wrong produce exactly the same symptom: requests keeping their old destination. That ambiguity is why this is worth diagnosing in a fixed order rather than by guesswork — most people start by editing the URL, which is the one variable that is usually correct.

In short: A base URL that is never read and one that is wrong produce identical symptoms. Authentication precedes routing, so a 401 localises to the credential and a 404 to the path.
You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment
Confirm the key works first. One command, no SDK, costs nothing:
curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.

Test the transport before the application

Before changing anything in your code, find out whether the endpoint works at all. One command per layer, in order — each one eliminates a class of causes, and a later test is meaningless if an earlier one failed.

# One command per layer. Run them in order -- each one eliminates a class
# of causes, and the order matters because a later test is meaningless if an
# earlier one fails.

BASE="https://aicomp.ai/v1"      # include the /v1 segment
KEY="$GATEWAY_API_KEY"

echo "--- 1. DNS: does the host resolve at all?"
getent hosts "$(echo "$BASE" | sed -E 's#https?://([^/]+).*#\1#')" || echo "DNS FAILS"

echo "--- 2. Transport: can we reach it, and who answers?"
curl -sS -o /dev/null -w "connect=%{time_connect}s tls=%{time_appconnect}s http=%{http_code}\n" \
     "$BASE/models" -H "Authorization: Bearer $KEY"

echo "--- 3. Routing: which host actually answered?"
curl -sS -D - -o /dev/null "$BASE/models" -H "Authorization: Bearer $KEY" \
  | grep -iE "^(HTTP/|server:|location:)" 

echo "--- 4. Auth: is the key the problem? 401 = key, 404 = path."
curl -sS -o /dev/null -w "with key:    %{http_code}\n" "$BASE/models" -H "Authorization: Bearer $KEY"
curl -sS -o /dev/null -w "without key: %{http_code}\n" "$BASE/models"

echo "--- 5. Model list: does this endpoint serve the ID you are sending?"
curl -sS "$BASE/models" -H "Authorization: Bearer $KEY" | grep -o '"id":"[^"]*"' | head -20

Step four is the one that saves the most time. Authentication happens before routing, so a 401 means the key was rejected and says nothing about your path, while a 404 means the path is wrong and the key is fine. Running both with and without the key takes four seconds and localises the fault.

The seven causes, most to least common

  1. A client that was built before the configuration changed. Clients read configuration at construction. A dev server that has been running since before you exported the variable keeps the old value for its whole lifetime.
  2. An override switch that was never flipped. Several tools keep their built-in provider unless you explicitly choose a custom one. The URL is correct and completely unused.
  3. The wrong path shape. Either a doubled /v1/v1 or a missing /v1. Both surface as a 404 that looks like a missing model.
  4. A stale credential in the environment. A long-lived shell holding the old key, or a .env that loads after your assignment and overwrites it.
  5. Model IDs validated against the original provider. Some clients check the model name against a known list before sending, so a valid gateway ID is rejected regardless of the URL.
  6. A proxy rewriting the destination. Nothing in the application mentions it, and every log line still names the host you configured.
  7. The wrong transport. Some CLIs require a specific API shape for custom providers and reject others at config load.

Print what the client holds, not what the config says

This is the check that resolves causes one, four and five in a single line. Whatever your configuration file states, the attribute on the client object is the truth.

# The most common cause of all: the client read the environment before you
# changed it. Print what the client actually holds -- not what your config says.

# OpenAI SDK
import os
from openai import OpenAI
c = OpenAI()
print("base_url :", c.base_url)
print("api_key  :", (c.api_key or "")[:6] + "...")   # prefix only, never log the whole key

# Any client that exposes the resolved URL: assert it at startup rather than
# discovering it from an error three services later.
assert str(c.base_url).rstrip("/").endswith("/v1"), "base URL did not take effect"

# Stale environment check -- the variable that actually wins depends on load order:
for k in ("OPENAI_BASE_URL", "OPENAI_API_BASE", "OPENAI_API_KEY"):
    print(f"{k:18s} = {(os.environ.get(k) or '<unset>')[:48]}")

The assertion is worth keeping. It turns a silent fallback into a startup failure, which is a far better place to discover it than three services downstream.

Tools that need a switch, not just a URL

Each of these has been verified against that tool's own documentation. The pattern is consistent: a correct base URL that is never consulted looks exactly like an incorrect one.

ToolThe switch people missWhat it looks like when missed
CursorOverride must be explicitly enabledThe client keeps using its own quota and never sends your request to the custom host.
ClineSelect an OpenAI Compatible providerChoosing a named provider validates model IDs against that provider's own list.
Roo CodeProvider plus model are configured togetherA model set without a matching provider falls back to the previous one.
AnythingLLMLLM_PROVIDER=generic-openaiWithout the provider selection the base path is read but never used.
Gemini CLIPick the auth type in /authUnset, it keeps prompting for a Google login instead of reporting a config error.
Codex CLIwire_api = "responses" for custom providersRecent versions reject a chat transport at config load, so the provider never starts.
AutoGenCapabilities must be declared explicitlyUndeclared capabilities are treated as absent, so tools silently stop working.
Agents SDKTracing is configured separately from the clientThe model request goes to your endpoint while traces still go to the platform.

Rule out a proxy before you change code

# Proxy variables are the cause people find last, because nothing in the
# application mentions them. A proxy can send your request to a different host
# entirely while every log line still names the endpoint you configured.

echo "HTTP_PROXY=$HTTP_PROXY"
echo "HTTPS_PROXY=$HTTPS_PROXY"
echo "NO_PROXY=$NO_PROXY"
echo "http_proxy=$http_proxy   (lowercase wins on some platforms)"
echo "https_proxy=$https_proxy"

# If a proxy is set and the gateway host is not excluded, requests are rewritten
# before they leave the process. Test by bypassing it entirely:
curl -sS --noproxy '*' -o /dev/null -w "direct: %{http_code}\n" "https://aicomp.ai/v1/models" \
  -H "Authorization: Bearer $GATEWAY_API_KEY"

# Permanent fix: exclude the host rather than unsetting the proxy globally.
#   export NO_PROXY="$NO_PROXY,$(echo 'https://aicomp.ai/v1' | sed -E 's#https?://([^/]+).*#\1#')" 

If bypassing the proxy changes the behaviour, exclude the host rather than unsetting the proxy globally — you almost certainly need it for everything else.

Why this matters in money, not just in principle

A request that silently goes to the default host is billed at the default host's rate. Nothing errors, nothing warns, and the usage appears in the account you were trying to move away from. Over a month of agent traffic, that is the difference between two invoices — and the failure is invisible until someone reconciles them.

The assert above is the cheapest guard against it: one line at startup, and the misconfiguration stops being silent.

What it costs

Models reachable once the endpoint actually takes effect — gateway rate vs official list
ModelGateway rate
in / out per 1M tokens
Official list
in / out per 1M tokens
Diff
gpt-5.6-luna$0.1 / $0.6$0.2 / $1.250%
claude-sonnet-5$1 / $5$2 / $1050%
deepseek-v4-flash$0.22 / $0.66— / ——

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget. Across the 216 models in our cleaned catalogue the median output-to-input ratio is 4.0×.

What actually goes over the wire

Every OpenAI-compatible call is an HTTP POST to a base URL plus a resource path, with two things that decide everything else: the path, and the bearer token.

POST /v1/chat/completions HTTP/1.1
Host: aicomp.ai
Authorization: Bearer sk-...
Content-Type: application/json

{"model": "gpt-5.6-luna",
 "messages": [{"role": "user", "content": "..."}],
 "stream": true}

Read that literally, because three separate failure modes hide in it. The Host comes from your base URL — if the request arrives somewhere unexpected, this is the line that was wrong. The Authorization header is the only thing identifying you, so a shared key means indistinguishable spend. And the model value inside the body is not validated against anything client-side: send an ID the server does not serve and you get a server-side error, not a client-side warning.

Nothing else in the request is provider-specific. The message array, the tool schemas, the stream flag and the response shape are the wire format, and they are what makes one base URL substitutable for another. This is also the reason the change is reversible: you are not recompiling anything, you are editing one string.

Symptoms, real causes, and how to tell them apart

SymptomWhat is actually happeningHow to confirmFix
Requests go to the old host with no errorThe client was constructed before the configuration changed, or an override switch was never enabled.Print base_url from the client object, not from your config file.Pass the base URL as an explicit argument and assert it at startup.
404 that mentions a missing modelThe path shape is wrong — a doubled or missing /v1 segment — rather than the model ID.curl /v1/models directly and compare the shape that returns a list.Match the shape that works in curl.
401 with a freshly created keyA stale key in a long-lived shell, or a .env that loads after your assignment.Print the first six characters of the key the client holds.Pass api_key explicitly rather than relying on the environment.
model_not_found for an ID the endpoint servesThe client validates model names against its own provider list before sending.Compare the ID against the endpoint's /v1/models output.Use the exact ID, and choose the custom provider rather than a named one.
Everything works in curl but not in the appA proxy rewriting the destination, or a client built at import time.Run with the proxy bypassed, and print the resolved URL after import.Exclude the host from the proxy; construct the client after configuration.
The CLI refuses to start with a custom providerSome CLIs require a specific transport for custom providers and reject others at config load.Read the startup error — it names the transport, not the URL.Declare the required transport for that provider.

When this is the wrong move

Three things that make this take longer

SituationWhy it breaksDo this instead
You are about to edit the URL a third timeThe URL is usually correct. The usual faults are that it is never read, or that the path shape is off by one segment.Print the resolved URL from the client object before changing anything.
You are debugging inside a long-running processThat process read its configuration once, at startup, and will not re-read it.Restart it, or pass routing as explicit arguments.
You plan to fix it by adding more environment variablesLoad order decides which variable wins, so adding one can make the ambiguity worse.Pass routing explicitly and assert it at startup.

Rolling it out without finding out the hard way

The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:

  1. Prove it on one request, with the resolved base URL asserted at startup. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
  2. Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
  3. Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
  4. Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.

Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.

FAQ

Requests still go to the default host. Is my URL wrong?

Not necessarily, and this is the distinction worth internalising: a URL that is never read and a URL that is wrong produce identical symptoms. Several tools keep their built-in provider unless you explicitly switch to a custom one, and a few validate model IDs against the original provider's list regardless of where you pointed them. Before you touch the URL, confirm the application is actually reading it — print the resolved base URL from the client object rather than from your configuration file.

Should the URL end in <code>/v1</code> or not?

It depends on what the client appends, and this is the single most common path mistake. Clients that follow the OpenAI convention append the resource path themselves, so the base URL should end at /v1 and no further. If you include the endpoint name, you get /v1/v1/… and a 404. If you omit /v1 entirely, the client appends to the wrong root and you get a 404 that looks like a missing model. The fastest way to settle it is one curl against /v1/models — whatever shape returns a list is the shape your client needs.

I get a 401 with a key I just created. What does that mean?

It means the key was rejected, and because authentication happens before routing, it also means your request never reached a model. The usual causes are a stale OPENAI_API_KEY in a long-lived shell, a .env that loads after your assignment and overwrites it, or passing the key to the wrong parameter name. Print the first six characters of the key the client actually holds — never the whole thing — and compare against what you expect.

What is the difference between a 401 and a 404 here?

The difference localises the fault, which is why it is worth learning. A 401 is the credential being rejected before routing happens, so it says nothing about your path. A 404 returned with a key that works elsewhere says the path is wrong and the key is fine. Mixing them up sends people to edit a base URL that was correct all along, or to rotate a key that was never the problem.

Could a proxy be redirecting my requests?

Yes, and it is the cause people find last because nothing in the application mentions it. Proxy environment variables can rewrite the destination outside your code entirely, so every log line still names the host you configured while the request goes somewhere else. Test with the proxy bypassed, and if that is the cause, exclude the host rather than unsetting the proxy globally — you probably need the proxy for everything else.

I set the environment variable but nothing changed.

Clients read configuration when the client object is constructed, not when a request is sent. A long-lived process, a dev server that has been running since before you exported the variable, or an import that builds the client at module load will all keep the old value for their entire lifetime. Passing the base URL as an explicit argument is the durable fix, because it does not depend on when anything was read.

Get API access