Kilo Code OpenAI Compatible Base URL: What Goes in the Field, and What Gets Rejected
You opened Settings → Providers → Custom provider in Kilo Code, set Provider API to OpenAI Compatible, pasted your base URL and key, and the model list populated by itself — so far so good. Then you tried to name the provider orcarouter/auto and Kilo refused it. Then you tried to select that same model and it was not in the list you just watched populate. Two rejections, two different rules, and nothing in the UI explains why they differ.
The short version: provider IDs and model IDs follow different rules in Kilo Code, and its Base URL field accepts a full endpoint URL — Cline's and OpenCode's take a prefix and append the path themselves. Get those two facts straight and the rest is four fields.
The four Kilo Code OpenAI Compatible fields
Open Settings → the Providers tab → scroll to the bottom → Custom provider. The dialog asks for Provider ID, Display name, Provider API, Base URL, API key, Models and optional Headers.
Set Provider API to OpenAI Compatible. The other two options are OpenAI Responses and Anthropic Messages, which are different wire protocols — picking the wrong one fails on message shape, not on auth, so it looks like a credential problem.
On the Base URL row, the docs are explicit: "When a valid URL is entered, Kilo automatically fetches available models from the endpoint if it exposes an OpenAI-compatible models endpoint." If that automatic fetch fails, there is a Fetch models button to trigger it on demand rather than re-entering the URL.
The Base URL field takes a full endpoint, not just a prefix
This is the part Kilo does differently. Every other client in this series treats the base URL as a prefix and appends the resource path itself. Kilo accepts all three of these:
Standard https://api.provider.com/v1
Full path https://api.provider.com/v1/chat/completions
Custom https://custom-endpoint.provider.com/api/v2/models/chat
That flexibility is useful when your endpoint is not shaped like OpenAI's, and dangerous in the ordinary case: paste the full-path form against a gateway and you can end up requesting /v1/chat/completions/chat/completions. Start with the standard form, https://aicomp.ai/v1. Move to the full-path form only if the standard one 404s, and never stack the two.
Provider IDs and model IDs are not the same kind of string
Provider ID allowed: lowercase letters, digits, hyphens, underscores
rejected: "/" and everything else
aicomp ✅
my_gateway-2 ✅
orcarouter/auto ❌ rejected
Model ID allowed: whatever your endpoint publishes, including "vendor/model"
claude-sonnet-5 ✅
vendor/model-name ✅
Kilo's docs state the provider rule as: "A unique identifier using lowercase letters, numbers, hyphens, or underscores (e.g., myprovider). This becomes the provider_id in the provider_id/model_id format." No slash, so a routing alias cannot be a provider ID.
The model half is a different object. Per the same docs: "provider_id is the key under provider; model_id is the key under provider.<provider_id>.models." Those keys carry whatever your endpoint publishes — slashes included. That asymmetry is the whole confusion.
Verify before you configure
export KILO_KEY=sk-your-gateway-key
curl -s https://aicomp.ai/v1/models \
-H "Authorization: Bearer $KILO_KEY" \
| jq -r '.data[].id' | head -40
IDs printed means the credential and the endpoint are both good. Invalid token means stop configuring and fix the key. Copy model strings from this output rather than typing them — Kilo's live list is fetched from the same place, but a typo in a manually added model produces model not found with no fallback.
https://aicomp.ai/v1).
Create one free →
What Kilo Code does differently on a custom endpoint
It pulls the model list live
Once the key and base URL are accepted, Kilo fetches the catalogue from your endpoint and shows it as a checklist. Under a generic provider in most other tools the list is a static snapshot shipped with the extension, and an empty dropdown tells you nothing. In Kilo, an empty list is real signal: it means the fetch failed, which is an auth or connectivity problem rather than a missing feature.
The caveat is that routing aliases do not appear in the pulled list. An alias like orcarouter/auto is a gateway-side shortcut, not a model row, so it is not in /v1/models and will never be offered. To use one, click + Add model and type it. The field accepts it, because model IDs allow slashes — it just will not be suggested to you.
Unset token limits disable context management entirely
This is the one Kilo-specific trap worth knowing before you configure anything. Kilo resolves limits in a fixed order: your config under provider.<id>.models.<model>.limit, then its built-in catalog snapshot, then a fallback. The fallback is not a safe default.
Per the docs: if your model is custom or local, is not in the built-in catalog, and you set no limits, both context and output resolve to 0. The consequences are concretely bad:
- Compaction is disabled. With
context: 0, Kilo skips overflow detection and the conversation grows unbounded until the provider rejects the request. - Output falls back to 32,000 tokens, which may be far above or far below what the model actually accepts.
- Context usage tracking is skipped, so the usage numbers you would use to compare models are missing.
When a model stops because it hit limit.output, Kilo shows a visible warning that the response may be incomplete. So it is not silent — but a truncated diff reads like the model failing at the task, and the warning is easy to scroll past.
Set both values explicitly. /v1/models publishes them for most endpoints:
curl -s https://aicomp.ai/v1/models -H "Authorization: Bearer $KILO_KEY" \
| jq -r '.data[] | select(.id=="claude-sonnet-5") | {context_length, max_output_tokens}'
The config file, in full
The VS Code dialog writes a config file, and the CLI reads one. Both are editable, and the schema is documented:
{
"$schema": "https://app.kilo.ai/config.json",
"model": "aicomp/claude-sonnet-5",
"provider": {
"aicomp": {
"options": {
"apiKey": "{env:AICOMP_API_KEY}",
"baseURL": "https://aicomp.ai/v1"
},
"models": {
"claude-sonnet-5": {
"name": "Claude Sonnet 5",
"tool_call": true,
"limit": { "context": 200000, "output": 16384 }
}
}
}
}
}
The key never goes in the file — {env:AICOMP_API_KEY} reads it from the environment. tool_call: true matters more than it looks: without it Kilo will not route tool use to the model, and a coding agent is nothing but tool calls.
The CLI reads ~/.config/kilo/kilo.jsonc and also accepts kilo.json; the VS Code extension writes kilo.json. If you configure one surface and debug the other you are editing a file nothing reads:
# Which one actually exists on your machine?
ls -la ~/.config/kilo/ 2>/dev/null
# JSONC allows comments; strip them before parsing with a strict JSON tool
sed 's|//.*||' ~/.config/kilo/kilo.jsonc | jq . > /dev/null && echo "parses clean"
# Confirm the provider is actually active, and list what Kilo resolved
kilo models
kilo models is the fastest way to check that your provider loaded at all — it lists available models and shows whether your provider is in the active set.
Azure OpenAI GPT-5 will not work through this provider
This one has a specific cause. Azure's GPT-5 rejects the max_tokens parameter that a generic OpenAI-compatible provider sends; it expects max_completion_tokens. The request fails on a parameter mismatch, not on auth or path, which is why it looks like every other failure in this article. The fix is not configuration — use Kilo's native azure provider instead of the generic one.
Failure modes
| Symptom | What is actually wrong | Fix |
|---|---|---|
401 immediately | Key pasted with a trailing space or newline, or the wrong credential for this endpoint | Re-paste using the copy button; confirm with curl before touching the panel |
404 on every request | Full-path base URL stacked on top of a client-appended path | Drop back to the standard https://host/v1 form |
404 despite a correct-looking URL | Endpoint genuinely not shaped like OpenAI's | Use the custom form and point at the exact endpoint path your gateway documents |
model not found | Manually added ID typed rather than copied | Copy the exact string from GET /v1/models |
| Wanted model missing from the pulled list | It is a routing alias, not a model row, so the fetch never returns it | Click + Add model and type it |
| Provider ID rejected | Slash in the provider ID — only letters, digits, hyphens and underscores are allowed | Rename to something like mygateway; keep the slash in the model ID |
| Response arrives incomplete | Model stopped at limit.output. Kilo shows a "response may be incomplete" warning, so it is not silent — just easy to miss | Raise limit.output for that model in the config file |
| Conversation grows until the provider rejects it | limit.context unset and the model is not in the built-in catalog, so compaction is disabled | Set limit.context explicitly. The fallback is 0, not a safe default |
| Azure GPT-5 fails on every call | Generic provider sends max_tokens; Azure GPT-5 requires max_completion_tokens | Switch to Kilo's native azure provider |
| Tasks succeed, gateway log empty | Wrong config surface — extension wrote kilo.json, you edited the CLI's kilo.jsonc | Check ~/.config/kilo/ and confirm the request in your usage log |
What a Kilo Code session costs
Kilo Code compacts conversation history to stay inside limit.context, and every compaction is a summarisation call you pay for. That gives the bill a third component most comparisons ignore: not just input and output per turn, but the periodic re-summarisation of everything that came before. Set limit.context too low and you pay for more compactions; leave it at 0 and you pay for unbounded context until the provider rejects the request outright.
One heavy day — roughly 30 turns with two compactions — lands near 160k input tokens and 70k output tokens. Across 20 working days, 3.2M input and 1.4M output.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-sonnet-5 | $1.00 / $5.00 | $0.51 | $10.20 |
| claude-opus-5 | $2.5 / $12.5 | $1.28 | $25.50 |
| gpt-5.6-sol | $2.5 / $15.00 | $1.45 | $29.00 |
| deepseek-v4-flash | $0.22 / $0.66 | $0.08 | $1.63 |
| gemini-3.7-flash | $0.375 / $1.875 | $0.19 | $3.83 |
| kimi-k3 | $1.5 / $7.5 | $0.77 | $15.30 |
One day = 160k input + 70k output on this page's workload. 20 days = 3.2M input and 1.4M output. Rates checked 2026-09-26 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
Because Kilo Code fetches the live list, comparing two rows on the same task is a picker selection — same key, same base URL, nothing else changes. One caution when you do it: a response that hit limit.output will make the cheaper model look worse than it is. Check for the incomplete-response warning before drawing any conclusion about capability.
qwen3.5-flash at $0.2 per 1M against gpt-realtime-2.1 at $12. Measured over the 216 cleaned models that publish an output rate. The extremes are wider still, but comparing a frontier reasoning model against a 9B one tells you nothing about a coding session — the percentile band is the range you actually choose within.FAQ
FAQ
Should the Kilo Code base URL end in /v1?
It should, but it is not required to. Kilo's Base URL field accepts the standard https://api.provider.com/v1, the full https://api.provider.com/v1/chat/completions, and arbitrary custom paths like https://custom-endpoint.provider.com/api/v2/models/chat. Start with the standard form. Only move to the others if the standard one 404s — stacking a full path on a client that appends its own produces a doubled path and a 404 on every call.
Why does Kilo reject my provider ID?
Because provider IDs allow only lowercase letters, digits, hyphens and underscores. A slash is rejected, so orcarouter/auto fails as a provider ID. Use a flat name like orcarouter or mygateway. Note that the same restriction does not apply to model IDs: those can be vendor/model, so the alias works fine as a model even though it fails as a provider.
Why is my model not in the list Kilo fetched?
Two possibilities. If the list is empty, the fetch failed — the docs say Kilo fetches automatically once a valid URL is entered, so an empty list means that fetch did not succeed. Click Fetch models to retry it on demand before you start editing the URL. If the list is populated but your model is absent, it is probably a routing alias rather than a model row, and /v1/models does not return aliases. Click + Add model and type the ID; the field accepts it.
Where does Kilo store the configuration?
The CLI reads ~/.config/kilo/kilo.jsonc and also accepts kilo.json. The VS Code extension writes kilo.json. Editing one while the other is the live file is a common source of "I changed it and nothing happened". Check what exists in ~/.config/kilo/ before debugging a setting that appears to be ignored.
Why does Azure GPT-5 fail through the OpenAI Compatible provider?
Because it is a parameter mismatch, not a credential or path problem. Azure's GPT-5 requires max_completion_tokens and rejects the max_tokens a generic OpenAI-compatible provider sends. No amount of base URL or key configuration fixes it — use Kilo's native azure provider for that deployment.
Do I need to fill in Model Configuration?
Yes, and specifically the limit object. Kilo resolves limits from your config first, then its catalog snapshot, then a fallback — and the fallback is 0. If your model is not in the built-in catalog and you set nothing, context and output both resolve to 0, which disables compaction entirely and drops output back to an internal 32,000-token default. Set limit.context and limit.output under provider.<id>.models.<model> to the model's real figures. /v1/models publishes both for most endpoints.
Two related flags are worth setting while you are in there: tool_call: true (without it Kilo will not route tool use to the model) and cost with input and output per million, if you want Kilo's own spend figure to mean anything.
Related
- Cline with an OpenAI-compatible endpoint — the same override in a sibling fork
- Base URL not working — doubled paths, wrong vendors, inert settings
- OpenAI-compatible error codes — what each status tells you and in what order
- Context length exceeded — when the error is context, not config
- OpenAI SDK with a custom endpoint — the raw-client version, useful for isolating Kilo from the endpoint
- All tools with one endpoint — the same four fields across editors and CLIs
- Model price index — compare output rates per model, because output is the side a Kilo Code session actually bills