Goose with a Custom OpenAI-Compatible Provider

In short: Goose reads GOOSE_PROVIDER, GOOSE_MODEL, OPENAI_HOST and OPENAI_API_KEY, and will not start without a model. The host variable is OPENAI_HOST, not the OPENAI_BASE_URL most tools use. In a custom provider file, base_url wants a version prefix and the models array supplies context sizes only — setting dynamic_models to false reduces Goose to exactly the IDs listed.

Goose is Block's open-source, on-machine agent: a terminal session that reads files, runs commands and extends itself through MCP servers. Unlike an editor plugin, it has no settings panel — everything is an environment variable or a file in ~/.config/goose/. That is why it is quick to point at a different endpoint and equally quick to point at the wrong one without noticing.

Three things make it less obvious than it looks. GOOSE_PROVIDER and GOOSE_MODEL are separate, and the second is mandatory — Goose will not start without a model, and the failure reads like a broken install. The host variable is OPENAI_HOST, not the OPENAI_BASE_URL most tools use, so a config copied from another client silently does nothing. And the models array in a custom provider file is not a model list — it supplies context sizes, and one flag turns it into a cage.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page covers both configuration routes, the two variable-name traps, what context_limit actually controls, how to verify the route without guessing, and what a month of autonomous agent work costs once you account for the fact that Goose re-sends the entire session on every turn.

Two routes, and when to use each

Route one: environment variables with the built-in provider. Fastest, and enough for one gateway.

export GOOSE_PROVIDER=openai
export GOOSE_MODEL=claude-sonnet-5
export OPENAI_HOST=https://your-gateway.example.com/v1
export OPENAI_API_KEY=your-key-here

Route two: a custom provider file. Needed when you want several endpoints selectable by name, or when the built-in provider's defaults fight your gateway.

{
  "name": "my_gateway",
  "engine": "openai",
  "display_name": "My Gateway",
  "api_key_env": "GATEWAY_API_KEY",
  "base_url": "https://your-gateway.example.com/v1",
  "models": [
    { "name": "claude-sonnet-5", "context_limit": 200000 }
  ]
}

Save it as ~/.config/goose/custom_providers/my_gateway.json, then select it:

export GATEWAY_API_KEY=your-key-here
export GOOSE_PROVIDER=my_gateway
export GOOSE_MODEL=claude-sonnet-5

Both routes persist through ~/.config/goose/config.yaml (macOS/Linux) or %APPDATA%\Block\goose\config\config.yaml (Windows):

GOOSE_PROVIDER: my_gateway
GOOSE_MODEL: claude-sonnet-5

A custom provider file shadows a built-in provider of the same name. Defining one called openai replaces the built-in rather than adding to it, which is occasionally what you want and usually not — name it something distinct.

The variable-name traps

Two, and both fail quietly.

OPENAI_HOST, not OPENAI_BASE_URL. Goose's built-in OpenAI provider reads the host from OPENAI_HOST. Most other tools on this site read OPENAI_BASE_URL or OPENAI_API_BASE, and integration docs written for those tools will hand you the wrong variable. Setting OPENAI_BASE_URL and leaving OPENAI_HOST unset does not produce an error about a missing variable — Goose falls back to the provider's default host and your requests go to the public API with your gateway key, which fails authentication in a way that looks like a key problem.

api_key_env names a variable; it is not a key. In a custom provider file, api_key_env = "GATEWAY_API_KEY" means "read $GATEWAY_API_KEY from the environment". Export it in the shell that launches Goose, or nothing is sent. GUI-launched terminals and editors that do not inherit shell exports are the usual cause of a key that works in one window and not another.

GOOSE_MODEL must be set. It is not optional and it is not inferred from the provider. Without it Goose refuses to start, and because the error appears at launch rather than at request time it is routinely misread as a broken installation.

What base_url wants

The same version-prefix ambiguity that appears in avante's endpoint field shows up here, and it is documented both ways depending on whose guide you read.

Published examples include a full path:

"base_url": "https://ai-gateway.vercel.sh/coding-agent/v1/chat/completions"

and a version prefix:

"base_url": "https://api.hpc-ai.com/inference/v1"

Goose appends the completions path, so the prefix form is what you want. Rather than trust either, read the log: send one prompt with goose run --text "say hello" and look at the recorded path on your gateway. A doubled /chat/completions means drop the suffix from base_url. Wrong host and wrong path produce the same error in the terminal and are only distinguishable server-side — the failure list is built around exactly that distinction.

The models array is not a model list

This is the most consequential misunderstanding in the file format, because it degrades silently over weeks rather than failing immediately.

models supplies context_limit values only. It tells Goose how much context a model can hold, which governs when Goose compresses or truncates the session. It does not enumerate what the model picker shows — Goose discovers the rest of the catalogue from the API.

Two consequences:

Leave dynamic_models alone. Setting it to false reduces Goose to exactly the models listed in the file. New models stop appearing, and a stale file quietly becomes a hard ceiling on what your agent can use. The value of the array is the context sizes, not the membership.

context_limit is a behavioural setting, not documentation. Get it wrong in either direction and the session degrades: too low and Goose compresses or truncates while plenty of headroom remains, losing earlier context mid-task; too high and it runs past what the model accepts and the request fails on length rather than being trimmed. If a long session loses its beginning or dies abruptly, this number is the first thing to check, and context-length errors covers the failure shapes.

Why the input column dominates here

Of every client on this site, a terminal agent running extended autonomous sessions has the largest input footprint, and the reason is structural rather than incidental.

Each turn sends the entire session history plus every tool result accumulated so far. A tool result is often a whole file. By turn thirty of a refactor, the request carries thirty turns of history and thirty file reads, and the model returns a few hundred tokens of reasoning plus a tool call. Then the next turn sends all of it again, plus that tool call.

A day of work on this page's workload is 260k input and 55k output tokens — 5.2M input and 1.1M output across 20 working days. That is the largest daily input figure on this site, and it is why the absolute numbers matter more here than the ratio.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-sonnet-5$1.00 / $5.00$0.54$10.70
gpt-5.6-sol$2.5 / $15.00$1.48$29.50
deepseek-v4-flash$0.22 / $0.66$0.09$1.87
glm-5.3$0.70 / $2.2$0.30$6.06
kimi-k3$1.5 / $7.5$0.80$16.05

One day = 260k input + 55k output on this page's workload. 20 days = 5.2M input and 1.1M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Across 219 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (150 models) charge at least 4× more for output than input. The other 624 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Compare the weighting against the editor clients. avante.nvim sits at 87% input because it returns a small applied diff against repository context; Goose sits lower on the ratio but far higher in absolute terms, because the context it re-sends is a whole session rather than a file selection. The practical difference is that session length is the cost lever here. Turning one long autonomous run into three shorter ones costs more in your attention and less in tokens, and no model choice moves the number as much.

For anything non-trivial, run the figures with your own split in the calculator rather than extrapolating from per-token rates — the ratio assumption does more damage at this scale than anywhere else.

Verifying the route

Cheapest first.

Does the endpoint respond at all? curl $GATEWAY/v1/models -H "Authorization: Bearer $KEY". A list of IDs validates host, path and credential together.

What does Goose think it is configured to use? goose info -v prints the resolved provider and model. If it names a provider you did not set, your environment variable is being overridden by config.yaml — the file wins over the shell in most setups.

Does a request actually arrive? One-shot it: goose run --text "say hello". Then check the gateway's request log for a timestamped entry. A successful curl from your machine proves connectivity from your machine; Goose's requests come from the same machine, so in this case it is meaningful — unlike an editor whose backend originates the call.

Is it using the key you think? The log entry will show the authenticated identity. A request that succeeds while your gateway dashboard shows no usage means it went somewhere else, which is the OPENAI_HOST trap above.

FAQ

FAQ

Why does Goose say no model is configured when I set a provider?

GOOSE_PROVIDER and GOOSE_MODEL are separate variables and both are required. Setting the provider alone leaves Goose with no model and it refuses to start, which is often misread as a broken installation.

Should I use `OPENAI_BASE_URL` or `OPENAI_HOST`?

OPENAI_HOST. Goose's built-in OpenAI provider reads that variable. OPENAI_BASE_URL is what most other tools use and has no effect here — leaving OPENAI_HOST unset sends requests to the provider's default host.

Should `base_url` include `/chat/completions`?

No — use the version prefix, https://host/v1. Goose appends the completions path itself. Published examples differ, so confirm from your gateway's request log rather than from whichever guide you found first.

What is the `models` array for if it does not list the models?

It supplies context_limit values, which govern when Goose compresses or truncates a session. Model discovery comes from the API. Do not set dynamic_models: false — that reduces Goose to exactly the listed models and stops new ones appearing.

Why did a long session lose its beginning or fail abruptly?

Almost always context_limit. Too low and Goose compresses or truncates while headroom remains; too high and the request exceeds what the model accepts and fails on length instead of being trimmed.

Why is Goose more expensive per session than an editor assistant?

Because every turn re-sends the full session history plus accumulated tool results, and tool results are often whole files. Editor clients send a file selection. Session length, not model choice, is the primary cost lever.

Related

Get API access