Codex CLI with a custom endpoint
Codex CLI is OpenAI's terminal agent. It reads two environment variables or one provider block, and once either is set every call bills against your own key. The part that trips people up in 2026 is not the URL — it is that Codex now requires the Responses wire format.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
What you need before you start
- An API key from the gateway, and the base URL (
https://aicomp.ai/v1) - Codex CLI installed (
npm install -g @openai/codex) - A model ID copied from the endpoint's own model list, not from memory
Option 1 — two environment variables
Enough for a test, and the form CI expects. Both values have to be set in the same shell that launches the CLI; a variable exported in one terminal does not reach a process started in another.
export OPENAI_API_KEY="sk-your-gateway-key"
export OPENAI_BASE_URL="https://aicomp.ai/v1"
codex "refactor this module to use the new client"
Option 2 — a provider block in config.toml
Better when you want the setting to persist, or when you want more than one route configured at once. The file lives at ~/.codex/config.toml.
# ~/.codex/config.toml
model_provider = "aicomp"
model = "gpt-5.6-luna"
[model_providers.aicomp]
name = "AICOMP Gateway"
base_url = "https://aicomp.ai/v1"
env_key = "AICOMP_API_KEY"
wire_api = "responses"
Three fields carry the weight. base_url decides where requests go, and it needs the /v1 segment because the CLI appends the rest of the path itself. env_key names the variable the credential is read from, which is what lets two providers hold two keys without re-exporting anything. And wire_api selects the protocol — see the next section, because this is the field that fails hardest.
The field that breaks: wire_api
Codex 0.122 and later removed the chat-completions wire format for custom providers. A config that sets wire_api = "chat" no longer loads at all — the CLI reports it as unsupported and exits before making a request. Custom providers must declare wire_api = "responses".
The consequence is a compatibility requirement that is easy to miss: the endpoint has to serve /v1/responses. Plenty of OpenAI-compatible endpoints implement /v1/chat/completions only, and against those, Codex will not work however the base URL is configured. Check that the responses route exists before spending time on anything else:
curl -s https://aicomp.ai/v1/models \
-H "Authorization: Bearer $OPENAI_API_KEY" | head -c 300
What it costs
| Model | Gateway rate in / out per 1M tokens | Official list in / out per 1M tokens | Diff |
|---|---|---|---|
| gpt-5.6-luna | $0.1 / $0.6 | $0.2 / $1.2 | 50% |
| gpt-6-astra | $5 / $25 | $10 / $50 | 50% |
| deepseek-v4-flash | $0.22 / $0.66 | — / — | — |
Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.
Codex sends the working context on every turn, so the bill tracks context size far more than it tracks the number of prompts you type. A session that touches twenty files resends all of them, repeatedly. Output is priced at a multiple of input on every model here — a median of 4.0× across the 216 we track — but on this shape the input side still dominates.
Codex CLI failures and what each one is really about
| Symptom | What is actually happening | How to confirm | Fix |
|---|---|---|---|
| Config refuses to load, mentions wire_api | Codex 0.122+ dropped the chat wire format for custom providers. | The error names wire_api explicitly and arrives before any request is made. | Set wire_api = "responses" and confirm the endpoint serves /v1/responses. |
| 401 with a key that works in curl | A stale OPENAI_API_KEY in the shell, or auth state cached from a previous login. | Print the variable in the same shell that launches Codex; test the key with curl separately. | Clear the stale variable, or remove the cached auth state and re-authenticate. |
| Stream disconnects before completion | The responses compatibility layer is not emitting the terminal event the CLI waits for. | Compare a raw responses call against a chat call to the same endpoint. | Use an endpoint that completes the responses stream properly, or fall back to a tool that speaks chat. |
| model_not_found on a name the endpoint serves | Aliases do not resolve through a gateway; concrete dated names do. | List /v1/models and copy the exact ID. | Pin the full model name in config.toml rather than a short alias. |
| 404 with a doubled /v1 in the path | The base URL already ends in /v1 and the CLI appended another. | Read the path in the error message. | Drop the suffix from the base URL. |
When this is the wrong move
Cases where pointing Codex at a gateway is the wrong move
| Situation | Why it breaks | Do this instead |
|---|---|---|
| Your endpoint only implements chat completions | Codex requires the responses route since 0.122; chat-only endpoints cannot serve it. | Use a tool that speaks chat completions, or an endpoint that exposes both routes. |
| You need a stable config across a team | Environment variables differ per shell and per machine; drift is silent. | Ship a config.toml with a named provider, or set the variables in the image. |
| You are about to run it unattended on a large repo | An agent loop that gets stuck keeps spending until something stops it. | Scope the task to a subdirectory and watch the usage log on the first run. |
Rolling it out without finding out the hard way
The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:
- Prove it on one read-only task on a single file. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
- Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
- Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
- Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.
Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.
FAQ
Why does Codex reject wire_api = "chat"?
Codex 0.122 and later removed the chat-completions wire format for custom providers. Loading a config that sets wire_api = "chat" fails outright with a message saying it is no longer supported. The practical consequence is narrower than it sounds but worth knowing before you start: the endpoint has to serve /v1/responses, not only /v1/chat/completions. If your provider exposes only the chat route, Codex will not work against it regardless of how the base URL is set.
Which takes precedence, the environment variable or config.toml?
Both are read, and the file is the one to reach for when you want the setting to survive a new shell. Environment variables are quicker for a one-off test and are the form containers and CI expect. Where the two disagree, check which one the CLI actually reports at startup rather than assuming — a stale OPENAI_API_KEY left over in your shell is the single most common cause of a 401 that looks like a broken base URL.
Why do I get model_not_found for a model the endpoint serves?
Codex sends the model string you give it verbatim, and gateways generally resolve concrete names rather than aliases. A short alias such as gpt-5-mini may not resolve where gpt-5-mini-2025-08-07 does. List the endpoint's models and copy the exact string instead of reconstructing it from memory.
What does env_key in the provider block mean?
It names the environment variable Codex reads the key from for that provider, so different providers can carry different credentials without you re-exporting anything. It is the mechanism behind keeping a work route and a personal route configured side by side.
Is the /v1 suffix on the base URL required?
Yes in practice. The CLI appends resource paths — /responses, /models — to whatever you give it, and a base URL without the version segment produces requests to the wrong path. A doubled /v1/v1/ is the other failure mode, and it is visible in the path the error returns.