avante.nvim with a Custom OpenAI-Compatible Endpoint

In short: avante is configured through a provider spec, not a URL field: inherit from openai with __inherited_from, give a version-prefix endpoint, and note that api_key_name names an environment variable. The endpoint is documented both as a full path and as a prefix — settle it from a request log. extra_request_body is the only route to max_tokens for deployments rejecting max_completion_tokens.

avante.nvim aims to be a Cursor-shaped experience inside Neovim: a sidebar, applied diffs, a file picker, repository context gathered automatically. That last feature is the one that decides your bill, and it is also the reason pointing avante at your own endpoint behaves differently from pointing a chat client at one.

The configuration itself is small. Two or three keys in a lazy.nvim opts table. What makes it fiddly is that the configuration is a provider spec, not a URL field, and three of its keys have behaviours that are not what the names suggest: endpoint carries a version-prefix ambiguity that has changed across releases, api_key_name is the name of a variable rather than a credential, and extra_request_body is the only way to reach models that reject a newer request field.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page walks the provider block, the version-prefix question and how to settle it from a log rather than a forum thread, the credential patterns, the max_completion_tokens problem, and what avante's automatic repository context does to a monthly bill.

The provider block

avante selects one provider by name and defines providers in a table:

{
  "yetone/avante.nvim",
  event = "VeryLazy",
  version = "*",
  opts = {
    provider = "my_gateway",
    providers = {
      my_gateway = {
        __inherited_from = "openai",
        endpoint = "https://your-gateway.example.com/v1",
        model = "claude-sonnet-5",
        api_key_name = "GATEWAY_API_KEY",
        extra_request_body = {
          temperature = 0.3,
        },
      },
    },
  },
  dependencies = { "nvim-lua/plenary.nvim", "MunifTanjim/nui.nvim",
                   "stevearc/dressing.nvim", "ibhagwan/fzf-lua" },
}

Three keys carry the weight.

__inherited_from = "openai" copies the built-in OpenAI provider's request and response handling — the SSE parsing, the message assembly, the tool-calling shape. Without it you would have to write parse_curl_args and parse_response yourself. For any gateway speaking the OpenAI protocol, inherit from openai rather than writing a provider from scratch.

provider must match a key in providers. A typo here does not error; avante falls back and you get a request to a provider you did not configure. If requests are arriving somewhere unexpected, check this string first.

api_key_name is the name of an environment variable, not the key. api_key_name = "GATEWAY_API_KEY" means "read $GATEWAY_API_KEY". It also accepts the cmd: prefix for secret managers — api_key_name = "cmd:op read op://personal/Gateway/key --no-newline" — which is the pattern to use for a version-controlled dotfile.

The /v1 question, and how to settle it

This is the ambiguity worth knowing about before you configure anything, because it appears in the project's own documentation in two forms.

The official AvanteProvider spec shows endpoint as a full path:

endpoint = "https://api.openai.com/v1/chat/completions"

While the openai-inheriting examples — OpenRouter, Groq, DeepSeek, Mistral — show it as a version prefix with no path:

-- OpenRouter
endpoint = "https://openrouter.ai/api/v1"
-- Groq
endpoint = "https://api.groq.com/openai/v1/"

Both appear in current documentation. The resolution is that avante appends the completions path itself, so the prefix form is what you want in current releases: https://host/v1. If you paste the full-path form into a recent build you get a request to .../chat/completions/chat/completions and a 404 that carries no hint about which half of the URL was wrong.

Do not take any of this on faith, including from this page — the behaviour has changed across releases and your build may differ. Settle it in thirty seconds with a log rather than a search:

  1. Set endpoint = "https://host/v1".
  2. Send one sidebar prompt.
  3. Look at the request log on your gateway or provider dashboard and read the path, not the status.
  4. If it shows a doubled path, drop the version: endpoint = "https://host".

The same check distinguishes the two failures that look identical from the editor side — a wrong host and a wrong path both surface as an error toast, and only the server-side log separates them. That distinction is the whole subject of the general base-URL failure list.

Note also that Ollama is a first-class provider and takes no /v1: endpoint = "http://127.0.0.1:11434". Do not inherit it from openai, and delete any legacy ollama entry in providers — having both defined is a known source of the endpoint being ignored.

max_tokens versus max_completion_tokens

One field, one failure mode, worth knowing cold because the symptom is a refusal rather than an error.

Newer OpenAI models expect max_completion_tokens. Plenty of OpenAI-compatible deployments — Mistral's is the documented example in avante's own provider list — expect the older max_tokens and reject the newer field. avante defaults to the newer one for its OpenAI-inherited providers.

The fix is one line in the provider spec:

extra_request_body = {
  max_tokens = 4096,
}

This is also the block for anything else the gateway wants in the body — temperature, top_p, presence_penalty — and it is per provider, so a gateway needing max_tokens does not force the setting onto providers that do not.

What the repository context costs

avante's headline feature is that you do not select context manually. It gathers it: the open buffers, the file picker selection, repository structure, and retrieved snippets from the project. That is the product. It is also the entire input column of your bill.

The asymmetry is sharper here than in almost any other client on this site. In a terminal agent like Goose the daily input is larger still — 260k tokens against this page's 200k — because each turn re-sends the whole session plus every tool result, and tool results are often whole files. In avante the request carries a chunk of your repository and the assistant returns an applied diff — usually tens of lines. Either way the input column decides the month. A day of work on this page's workload is 200k input and 30k output tokens — 4M input and 0.6M output across 20 working days.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-sonnet-5$1.00 / $5.00$0.35$7.00
gpt-5.6-sol$2.5 / $15.00$0.95$19.00
deepseek-v4-flash$0.22 / $0.66$0.06$1.28
glm-5.3$0.70 / $2.2$0.21$4.12
kimi-k3$1.5 / $7.5$0.53$10.50

One day = 200k input + 30k output on this page's workload. 20 days = 4M input and 0.6M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Across 219 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (150 models) charge at least 4× more for output than input. The other 624 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Read that weighting before you optimise anything, because it inverts the usual advice. Choosing a model by its output rate — the reflex most people bring from chat pricing — optimises 13% of this bill and ignores the other 87%. The calculator takes your own split rather than a default, which matters here more than usual: a workflow that uses the sidebar for whole-file generation lands on a different ratio than one that uses it for review.

Two things follow for cost control specifically. The file picker is the expensive control, not the model picker — every added file is input tokens on every turn. And long sessions compound, because context accumulates across turns; the same effect that produces context-length errors in other clients shows up here as a rising per-turn cost long before it shows up as a failure.

Before you debug the endpoint

Three checks, ordered cheapest first, because two of them eliminate the most common causes in under a minute.

Did the build step run? avante compiles a native component. With lazy.nvim pinned to a release this is handled; installing from source requires make (or make BUILD_FROM_SOURCE=true). A plugin that loads and then fails on first request, with no endpoint error, is very often an unbuilt binary rather than a network problem.

Are the dependencies present? Seven are required, four of them structural: plenary.nvim, nui.nvim, dressing.nvim and a picker such as fzf-lua. A missing one surfaces as a startup error naming the module, which is sometimes mistaken for a config error.

Does the key resolve? api_key_name names a variable; if that variable is unset in the shell that launched Neovim, the request goes out with an empty credential. GUI-launched editors are a common cause here, because they do not inherit a shell's exported variables.

FAQ

FAQ

Should `endpoint` include `/v1` or the full `/chat/completions` path?

Use the version prefix, https://host/v1, on current builds — avante appends the completions path itself. The full-path form appears in the provider spec documentation but produces a doubled path on recent releases. Confirm from your gateway's request log, which shows the resolved path.

Why is my OpenAI-compatible model rejected before it answers?

Usually a request-body field. Newer OpenAI models expect max_completion_tokens while several compatible deployments — Mistral is the documented case — expect max_tokens. Add extra_request_body = { max_tokens = 4096 } for the provider that needs it.

Can I put the API key in the config file?

Not directly — api_key_name names an environment variable holding the key. Use the cmd: prefix to read from a secret manager if the dotfile is version controlled, and note that the command runs on each provider resolution.

Does avante work with Ollama through the same setting?

Ollama is a first-class provider and its endpoint takes no version suffix: endpoint = "http://127.0.0.1:11434". Do not inherit it from openai, and remove any legacy ollama entry from providers.

Why is my bill higher than the chat volume suggests?

Because avante gathers repository context automatically — open buffers, picked files, retrieved snippets — and sends it each turn, while the returned diff is small. Input dominates. Reducing the files you add to context reduces cost more than switching to a cheaper output rate.

Do I need a custom provider for a gateway?

Not necessarily. Defining a provider with __inherited_from = "openai" plus endpoint and api_key_name is enough for any OpenAI-compatible gateway. Write parse_curl_args and parse_response by hand only for a genuinely non-conforming API.

Related

Get API access