CodeCompanion.nvim with an OpenAI-Compatible Endpoint
You added an openai_compatible adapter, opened a chat buffer, and got an answer — so the endpoint works. Then you selected three lines and hit the inline prompt, and it went somewhere else entirely. Or your key is in 1Password and you do not want it in a dotfile that gets committed. Or the chat worked yesterday and today every request fails with a connection error you cannot attribute to anything you changed.
CodeCompanion is the Neovim plugin where those three failures share one cause: an adapter is not a global endpoint setting. It is a named object that each strategy resolves independently, and the shipped openai_compatible adapter carries a default URL that is not OpenAI's. Once you see the configuration that way — four strategies, each with its own adapter reference, on top of an adapter whose default points at your own machine — the behaviour stops being mysterious.
https://aicomp.ai/v1).
Create one free →
This page covers the adapter block, the two schema shapes you will meet in the wild, why the default URL matters more than it looks, how to keep a key out of the file, which strategies silently fall back, and what a month of Neovim-native assistance actually costs once you account for the buffer contents the inline strategy sends whether you asked it to or not.
What an adapter is, and what it is not
CodeCompanion talks to models through adapters. An adapter is a Lua table describing a protocol: which URL to hit, how to authenticate, how to map the request schema, how to parse the response stream. Built-in adapters cover OpenAI, Anthropic, Copilot, Azure, Gemini, Ollama and the generic openai_compatible.
The thing that catches people is the second half. Configuring one adapter does not tell CodeCompanion to use it. Selection happens per strategy, and CodeCompanion ships four:
chat— the chat buffer, opened with:CodeCompanionChatinline— the prompt applied to a selection or the current buffercmd— the command-line strategy used by the prompt librarybackground— background tasks the agent runs without a visible buffer
Each takes adapter = "name". Leave one unset and it falls back to the default, which is openai. So a config that sets adapters.http.my_provider and only points strategies.chat at it will have a working chat buffer and an inline strategy still calling OpenAI with a key you did not intend to use. This is the single most common half-configured state, and it produces no error — the unconfigured strategy simply works, on the wrong provider.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-sonnet-5 | $1.00 / $5.00 | $0.32 | $6.30 |
| gpt-5.6-sol | $2.5 / $15.00 | $0.88 | $17.50 |
| deepseek-v4-flash | $0.22 / $0.66 | $0.05 | $1.08 |
| glm-5.3 | $0.70 / $2.2 | $0.18 | $3.50 |
| kimi-k3 | $1.5 / $7.5 | $0.47 | $9.45 |
One day = 140k input + 35k output on this page's workload. 20 days = 2.8M input and 0.7M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
The default URL nobody expects
Here is the detail that generates the most confusing support threads, and it is worth quoting precisely. The openai_compatible adapter's url field is optional, and when you omit it the default is http://127.0.0.1:11434 — the Ollama port.
That default exists because the adapter was introduced primarily for local models. It is a reasonable default for that use case and a trap for everyone else:
require("codecompanion").setup({
adapters = {
my_gateway = function()
return require("codecompanion.adapters").extend("openai_compatible", {
env = {
url = "https://your-gateway.example.com/v1",
api_key = "GATEWAY_API_KEY",
chat_url = "/v1/chat/completions",
},
})
end,
},
strategies = {
chat = { adapter = "my_gateway" },
inline = { adapter = "my_gateway" },
},
})
Two fields in that block are worth understanding separately from the URL.
url is the host and version prefix. https://host/v1, with no trailing slash and no path. The adapter appends chat_url to it.
chat_url is the path, and it defaults to /v1/chat/completions. Note what that means if your url already carries /v1: the default chat_url produces https://host/v1/v1/chat/completions. Set chat_url = "/chat/completions" in that case, or drop the version from url. Whether a given deployment wants the version in the host or the path is exactly the kind of thing that separates a working request from a 404, and the failure looks identical in the chat buffer either way.
api_key is the name of an environment variable, not the key. api_key = "GATEWAY_API_KEY" means "read the key from $GATEWAY_API_KEY". Putting the literal key there will not work, and this is the field where the distinction costs the most time because the error surfaced is an auth failure rather than a config error.
Two schema shapes, both real
Search for CodeCompanion configurations and you will find the adapter block in two different places. Both work; which one you need depends on the version you installed.
The flat form puts adapters directly under adapters:
adapters = {
my_gateway = function() ... end,
}
The namespaced form puts them under adapters.http:
adapters = {
http = {
my_gateway = function() ... end,
},
}
The namespaced form exists because CodeCompanion now supports transports other than HTTP — the ACP-style agent adapter being the obvious one — and the namespace leaves room. The flat form still resolves in current versions because that is where the built-ins live.
If your adapter seems to be ignored entirely, this is the second thing to check after strategy selection. An adapter defined under adapters.http while your config references it as a flat name, or the reverse, resolves to nothing and the strategy falls back to the default — again silently, again with a working chat buffer on the wrong provider.
Keeping the key out of the file
The cmd: prefix lets api_key name a command whose stdout becomes the key:
env = {
api_key = "cmd:op read op://personal/Gateway/credential --no-newline",
}
This is the right pattern for a dotfile under version control, and it is worth using even when the key is not especially secret, because the alternative is a plaintext credential in a repository that will eventually be public.
Two constraints. The command runs on every adapter resolution, so something slow — a hardware-key-backed op invocation, a network round trip — adds latency to the first request of every session. And it must print the key with no trailing newline, hence --no-newline; a key with an embedded \n fails authentication in a way that looks like a wrong key rather than a formatting bug.
What the inline strategy sends
The cost asymmetry on this page is not the one most people assume, and it comes from how inline works.
A chat buffer sends what you typed plus whatever variables and slash commands you referenced. The inline strategy sends the buffer, or the selection, plus surrounding context, every time. On a 4,000-line file with a three-line selection, the selection is not the request — the request is the selection embedded in a chunk of the file with an instruction wrapper.
That inverts the usual weighting for a coding tool. In terminal agents like Claude Code or OpenCode, the output column dominates because the agent writes files. In a Neovim inline workflow the assistant edits in place, the diff it returns is small, and the input column carries the month. A day of work on this page's workload is 140k input and 35k output tokens — 2.8M input and 0.7M output across 20 working days.
That ratio is why the percentile column in the table above matters more than the absolute rates here: a model that is 30% more expensive per output token is a rounding error on this workload, while the same percentage on input is not. Our guide to the cheapest models by workload shape starts from the split rather than the tool, which is usually the cheaper way to pick.
Verifying before you blame the model
Cheapest test first, because two of these cost ten seconds and one costs an afternoon.
Confirm the endpoint speaks the protocol at all. curl $URL/v1/models with your key. A list of IDs means the host, path and credential are all good. Anything else localises the problem to the URL, not the model — and the error-code map is the fastest way to read what comes back.
Confirm CodeCompanion is using the adapter you think. In a chat buffer, gd opens the buffer's debug contents including the resolved settings, and ga switches the adapter for that buffer. If gd shows the built-in openai schema, your strategy configuration is not taking effect.
Confirm the model supports tool calling. CodeCompanion's agent behaviour — file reads, edits, terminal commands — runs through tool calls, and tool call formats differ per model family. Models that do not emit them in the expected shape will chat fluently and then fail the moment a tool is needed. This is observed behaviour reported by users rather than a documented support matrix, so treat it as "test before relying on it" rather than a fixed list.
Confirm treesitter parsers are installed. :TSInstall markdown markdown_inline. The chat buffer is markdown-rendered, and missing parsers produce rendering breakage that looks like a plugin bug.
When the plugin is the wrong layer
Two situations where configuring CodeCompanion harder is wasted effort.
You want one endpoint for every tool, not just the editor. CodeCompanion is one adapter of four strategies among many clients. A proxy that every tool points at — including the ones without a plugin — is a different and usually simpler answer; the tooling matrix covers that shape.
You need per-request cost attribution across a team. A per-editor adapter gives you per-editor configuration, which is the opposite of centralised. Observability belongs at the gateway, which is what monitoring API spend is about.
FAQ
FAQ
Why does CodeCompanion try to connect to 127.0.0.1:11434?
Because the openai_compatible adapter's url field defaults to the Ollama port when you leave it unset. The adapter was introduced for local models and kept that default. Set env.url to your gateway's host and version prefix explicitly.
Chat works but the inline prompt uses the wrong provider. Why?
Because adapter selection is per strategy. Setting strategies.chat.adapter does not change inline, cmd or background; each falls back to the default openai adapter when unset. Point all four at your adapter name.
Should `url` include `/v1`?
It depends on what you put in chat_url, which defaults to /v1/chat/completions. If url is https://host/v1, set chat_url = "/chat/completions" — otherwise the two concatenate into /v1/v1/chat/completions. Check the resolved path in the debug buffer rather than guessing.
Can I put the API key directly in the config?
No — api_key names an environment variable holding the key, it does not hold the key. For secrets in version-controlled dotfiles use the cmd: prefix to read from a secret manager, and remember the --no-newline flag so the key is not authenticated with a trailing newline.
Why do tool-using prompts fail on some models?
CodeCompanion's agent behaviour depends on the model emitting structured tool calls, and tool call formats differ across model families. A model can produce fluent prose and still fail here. Test tool behaviour explicitly rather than inferring it from a successful chat reply.
Does using the inline strategy cost more than the chat buffer?
Per request, yes in input terms: inline sends buffer or selection context on every invocation, while chat sends what you reference. Because the returned diff is small, the input column dominates the monthly bill — the opposite weighting from a terminal agent that writes whole files.
Related
- avante.nvim with a custom endpoint — the other Neovim assistant, configured through providers rather than adapters
- Goose with a custom provider — the terminal-agent shape, where each turn re-sends the whole session
- Zed with an OpenAI-compatible endpoint — the editor that keeps two separate endpoint settings
- Every tool with one endpoint — the full matrix of base-URL field names