Crush with a Custom OpenAI-Compatible Provider
Crush is Charm's terminal agent — a TUI that reads your files, runs commands, talks to LSP servers and extends itself through MCP servers. It has no settings panel to speak of. Every provider lives in a single JSON file, and the switcher you reach with Ctrl+L or /model reads that file and nothing else.
That makes it fast to point somewhere new and easy to get subtly wrong. The type field is spelled openai in Crush's own README and openai-compat in a vendor integration doc published alongside it — both are official, and only one may be accepted by your build. Nothing is fetched from your endpoint: the models array is you declaring what exists, including its context window, which means a wrong number there silently changes how your sessions behave. And api_key holds the name of an environment variable, not a key — paste a real key and it works, then sits in your dotfiles forever.
https://aicomp.ai/v1).
Create one free →
This page covers where the file goes, how the provider block is assembled, how to settle the type and /v1 questions from evidence instead of guessing, what context_window and default_max_tokens actually control, and what a month of agent work costs once you account for Crush re-sending the session on every turn.
Where the file lives
# Linux / macOS
~/.config/crush/crush.json
# Windows
%USERPROFILE%\.config\crush\crush.json
The file does not exist until you create it. Crush starts fine without one — it just has no custom providers.
The provider block
{
"$schema": "https://charm.land/crush.json",
"providers": {
"mygateway": {
"type": "openai",
"base_url": "https://your-gateway.example.com/v1",
"api_key": "$GATEWAY_API_KEY",
"models": [
{
"id": "claude-sonnet-5",
"name": "Claude Sonnet 5",
"context_window": 200000,
"default_max_tokens": 32768,
"can_reason": true
}
]
}
}
}
Then:
export GATEWAY_API_KEY=your-key-here
cd /path/to/project
crush
type is not settled — resolve it from your own build
This is the field most likely to cost you an afternoon. Across official sources:
- Crush's own README configures Ollama and LM Studio with
"type": "openai", and an Anthropic-compatible endpoint with"type": "anthropic". - A vendor's published integration doc for the same client uses
"type": "openai-compat".
Both are describing the same client. Which one your build accepts is a function of the version you installed, not of which page you read.
Do not resolve this by reading more tutorials. Resolve it from the client:
crush --version
and then check the schema the file itself points at — "$schema": "https://charm.land/crush.json". If a provider you defined does not appear in the switcher at all, the type value is the first suspect, ahead of the URL.
api_key names a variable
The value is "$GATEWAY_API_KEY", a reference. Crush expands it at request time. Two consequences:
- A literal key works, so the mistake is invisible — it works right up until you commit the file or share a dotfiles repo.
- The variable must be exported in the shell that launches Crush. A key set in one terminal, or in a profile that only loads for login shells, produces a 401 that reads like a bad key rather than a missing one.
Whether base_url ends in /v1
Crush appends the chat completions path to whatever you give it, which is why the trailing segment matters and why sources disagree:
- A gateway integration guide states the base URL must end with
/v1. - A vendor integration doc for Crush shows a host with no path at all.
- The README's local-model examples end in
/v1/.
Whether you need it depends on whether your gateway expects the version prefix. There is exactly one reliable test — read the request path your gateway actually received:
# gateway log / dashboard, most recent request
POST /v1/v1/chat/completions ← you included /v1 and the gateway adds it too
POST /chat/completions ← the gateway forwards the versioned path already
Fix it in whichever direction produces a single /v1. Guessing produces a 404 that looks like a dead endpoint rather than a doubled path.
The models array declares context, it does not discover models
Crush does not call /v1/models on your endpoint. Everything in the picker comes from this array, which means:
- A model you do not list does not exist as far as Crush is concerned, even if your gateway serves it.
context_windowis a declaration, not a measurement. Too small and Crush starts trimming or summarising a session that your endpoint could have carried whole. Too large and you get context errors from the server on requests Crush believed were fine.default_max_tokenscaps each response. A patch or file write longer than this comes back truncated, and the failure looks like a model that "gives up halfway" rather than a configuration value.can_reasonaffects which tasks the model is offered for, not whether it can reason.
Why the input column dominates here
Crush is a terminal agent, and terminal agents pay for context twice: the request carries the session so far, and every tool result that came back during it. A file Crush read four turns ago is still in the request now. The output — a diff, a patch, a short explanation — is usually the smaller half.
A day of work on this page's workload is 250k input and 50k output tokens — 5M input and 1M output across 20 working days. On a model where output is billed at a multiple of input, that split is what decides whether the month is affordable, which is why the input rate is the number to negotiate and the output multiple is the number to model.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-sonnet-5 | $1.00 / $5.00 | $0.50 | $10.00 |
| gpt-5.6-sol | $2.5 / $15.00 | $1.38 | $27.50 |
| deepseek-v4-flash | $0.22 / $0.66 | $0.09 | $1.76 |
| glm-5.3 | $0.70 / $2.2 | $0.29 | $5.70 |
| kimi-k3 | $1.5 / $7.5 | $0.75 | $15.00 |
One day = 250k input + 50k output on this page's workload. 20 days = 5M input and 1M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
Verifying the route
crush --version, then confirm the provider appears underCtrl+L//model. Missing entirely =typeor JSON syntax.- Pick the model and send a one-line message. A reply proves routing, auth, streaming and the model id in one step.
- Check the gateway log for a timestamped request. Absent = wrong host or a
/v1mismatch; present with a 401 = unexported variable. - Ask it to read a large file and then reference it several turns later. That is the test the
context_windowvalue has to survive.
FAQ
FAQ
Why does my provider not appear in the model switcher at all?
The type value or the JSON itself. Check the file parses (jq . ~/.config/crush/crush.json), confirm the path matches your OS, and try the other spelling of type. A provider Crush cannot parse is dropped silently rather than reported.
Should `base_url` end in `/v1`?
It depends on whether your gateway forwards the version prefix, and official sources give both forms. Read the request path in your gateway log: one /v1, not zero and not two.
Can I put the API key directly in the file?
Yes, and that is the problem — it works, and then it is in your dotfiles. Use "$VAR" and export it in the shell that starts Crush.
Why does Crush truncate a long file write?
default_max_tokens on that model entry caps a single response. Raise it, but keep it below what the endpoint will actually return.
Why does a long session lose its beginning?
context_window in your declaration. If you declared less than the endpoint supports, Crush starts managing context earlier than it needs to.
Is Crush cheaper than an editor assistant per task?
Not per token, and the shape is what differs. Terminal agents re-send session history and accumulated tool results on every turn, so their input volume is among the highest on this site — see the Goose workload for the same pattern, and why output pricing dominates for what the multiple does to the bill.
Related
- OpenCode with a custom provider — another terminal agent, but it separates the credential from the config file
- CodeCompanion.nvim — the editor-side version of the same problem, where the endpoint is per-strategy rather than global
- Tools that take one endpoint — the full matrix
- Official list prices vs gateway rates — which number on this page is which