Open Interpreter with a Custom OpenAI-Compatible Endpoint
Open Interpreter runs a language model against a real shell on your machine. It writes code, asks before executing it, and feeds the output back to the model. Underneath, it talks to models through LiteLLM, which is why the configuration has three fields that most clients do not have and one field you cannot leave out even when it is meaningless.
--api_key is required even when your server ignores it. LiteLLM refuses to send a request without one, so the documented value for a local server is the literal string fake_key. --api_base does not carry any model metadata with it — context_window and max_tokens stay at their defaults, and in local mode those defaults are 3000 and 1000. And in Python the model string needs a openai/ prefix, because LiteLLM uses that prefix to decide which wire format to send.
https://aicomp.ai/v1).
Create one free →
This page covers the command line and the Python object, why the context numbers are the real failure mode, and what a month of interpreted sessions costs once every shell output is fed back in.
The command line
interpreter --api_base "https://gw.example.com/v1" --api_key "fake_key"
That is the whole connection. The --api_key value is not checked by your gateway in this case — it exists because LiteLLM requires the field to be non-empty. Leave it out and you get an authentication error that looks like your gateway rejected you, when in fact nothing was sent.
For a remote gateway with a real key, put the real key there.
The Python object
The Python package exposes the same three fields as attributes:
from interpreter import interpreter
interpreter.offline = True # disables online features
interpreter.llm.model = "openai/gpt-6-sol" # openai/ prefix selects the wire format
interpreter.llm.api_key = "fake_key" # LiteLLM requires a non-empty value
interpreter.llm.api_base = "https://gw.example.com/v1" # any OpenAI-compatible server
interpreter.chat()
The openai/ prefix is the part that trips people up. LiteLLM infers the provider from the model string, so "gpt-6-sol" on its own is ambiguous and LiteLLM may pick a different provider and a different request shape. Prefixing with openai/ forces the OpenAI format, which is what an OpenAI-compatible gateway expects.
interpreter.offline = True is separate from the endpoint. It turns off the online features — the hosted procedure library — so that a session you intend to be entirely on your gateway does not also reach out to a hosted service.
The context numbers do not come from the endpoint
This is the failure that costs the most time to diagnose, because nothing errors.
In local mode, Open Interpreter sets context_window to 3000 and max_tokens to 1000. Those are numbers suited to a small local model, and they are not inferred when you supply --api_base — so pointing at a remote gateway that serves a 200k-context model leaves you with a 3000-token window unless you set it:
interpreter --api_base "https://gw.example.com/v1" \
--api_key "fake_key" \
--model openai/gpt-6-sol \
--context_window 128000 \
--max_tokens 8192
The constraint to respect is max_tokens smaller than context_window. What you see when you get it wrong is not an error but truncation: long sessions lose their beginning, and long code output gets cut off mid-function. The model then "fixes" a truncated file it never saw the end of.
If responses are slow or fail, the documented move is to try a much shorter window first, around 1000, because smaller windows use less memory — then raise it once you know the endpoint is healthy.
Profiles instead of repeating flags
Rather than typing the same four flags every session, Open Interpreter reads YAML profiles:
interpreter --profiles # opens the profiles directory
interpreter --profile my_gateway.yaml
default.yaml is what loads when you do not name one, so putting your endpoint, model and context numbers there makes every session start connected. This is also the only clean way to keep two endpoints — a local one for experiments and a remote one for real work — without remembering which flags go with which.
Why the input side grows faster here
An interpreted session has a shape no chat client has: the model writes code, the shell runs it, and the output of that execution is fed back as input on the next turn. A command that prints two thousand lines of stack trace is billed as input. The loop repeats until the task is done.
A day of work on this page's workload is 260k input and 35k output tokens — 5.2M input and 0.7M output across 20 working days. The code the model writes is small; the terminal output it has to read back is not. See output vs input pricing for how to tell which side your own workload actually sits on.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| gpt-6-astra | $5.00 / $25.00 | $2.18 | $43.50 |
| deepseek-v4-flash | $0.22 / $0.66 | $0.08 | $1.61 |
| glm-5.3 | $0.70 / $2.2 | $0.26 | $5.18 |
| MiniMax-M3 | $0.15 / $0.60 | $0.06 | $1.20 |
| qwen3.8-max | $1.00 / $3.00 | $0.36 | $7.30 |
One day = 260k input + 35k output on this page's workload. 20 days = 5.2M input and 0.7M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
Verifying the route
- Send
--api_keyeven for a keyless server.fake_keyis the documented placeholder; an omitted key fails before anything reaches your gateway. - Prefix the model with
openai/in Python. Without the prefix LiteLLM guesses the provider, and the guess changes the request format. - Set
--context_windowand--max_tokensexplicitly. Nothing about--api_basetells Open Interpreter how large your model's window is, and the defaults are 3000 and 1000. - Run
--verboseonce. It prints the request details, which is the fastest way to confirm the host, path and model id that actually went out. - Ask for a long output on purpose — a directory listing of a large folder. If it comes back truncated, the context numbers are still at their defaults, not the endpoint.
FAQ
FAQ
Why is api_key required when my local server does not check it?
LiteLLM, which Open Interpreter uses to reach models, requires a non-empty key. Use fake_key, or your real key against a gateway that checks.
Do I need the openai/ prefix on the model name?
In Python, yes. LiteLLM infers the provider from the model string, and the prefix is what selects the OpenAI request format your gateway expects.
My responses are truncated but there is no error
context_window and max_tokens are still at their defaults — 3000 and 1000 in local mode. They are not inferred from --api_base. Set both, and keep max_tokens below context_window.
What does interpreter.offline do?
It disables the online features such as the hosted procedure library, so the session does not make calls beyond the endpoint you configured. It is independent of api_base.
Can I keep more than one endpoint?
Yes — save each as a YAML profile with interpreter --profiles and select with --profile. default.yaml is used when you do not name one.
How is this different from LiteLLM configuration?
Open Interpreter uses LiteLLM internally, so the provider prefixes and the key requirement come from there, but you set them on interpreter.llm rather than in a LiteLLM config. For LiteLLM's own surface see LiteLLM with a custom endpoint, and for a terminal agent with a review sandbox instead of live execution, Aider.
Related
- LiteLLM — the library Open Interpreter uses to reach models
- Aider — terminal agent that patches files rather than executing code
- Plandex — long-running agent with a review sandbox
- OpenAI-compatible error codes — what a 401 or 404 from your gateway actually means
- Tools that take one endpoint — the full matrix