Open Interpreter with a Custom OpenAI-Compatible Endpoint

In short: api_key is required even when nothing checks it, because Open Interpreter reaches models through LiteLLM — fake_key is the documented placeholder. api_base carries no model metadata, so context_window and max_tokens stay at 3000 and 1000 unless you set them, and in Python the model string needs an openai/ prefix to select the request format.

Open Interpreter runs a language model against a real shell on your machine. It writes code, asks before executing it, and feeds the output back to the model. Underneath, it talks to models through LiteLLM, which is why the configuration has three fields that most clients do not have and one field you cannot leave out even when it is meaningless.

--api_key is required even when your server ignores it. LiteLLM refuses to send a request without one, so the documented value for a local server is the literal string fake_key. --api_base does not carry any model metadata with it — context_window and max_tokens stay at their defaults, and in local mode those defaults are 3000 and 1000. And in Python the model string needs a openai/ prefix, because LiteLLM uses that prefix to decide which wire format to send.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page covers the command line and the Python object, why the context numbers are the real failure mode, and what a month of interpreted sessions costs once every shell output is fed back in.

The command line

interpreter --api_base "https://gw.example.com/v1" --api_key "fake_key"

That is the whole connection. The --api_key value is not checked by your gateway in this case — it exists because LiteLLM requires the field to be non-empty. Leave it out and you get an authentication error that looks like your gateway rejected you, when in fact nothing was sent.

For a remote gateway with a real key, put the real key there.

The Python object

The Python package exposes the same three fields as attributes:

from interpreter import interpreter

interpreter.offline = True                              # disables online features
interpreter.llm.model = "openai/gpt-6-sol"              # openai/ prefix selects the wire format
interpreter.llm.api_key = "fake_key"                    # LiteLLM requires a non-empty value
interpreter.llm.api_base = "https://gw.example.com/v1"  # any OpenAI-compatible server

interpreter.chat()

The openai/ prefix is the part that trips people up. LiteLLM infers the provider from the model string, so "gpt-6-sol" on its own is ambiguous and LiteLLM may pick a different provider and a different request shape. Prefixing with openai/ forces the OpenAI format, which is what an OpenAI-compatible gateway expects.

interpreter.offline = True is separate from the endpoint. It turns off the online features — the hosted procedure library — so that a session you intend to be entirely on your gateway does not also reach out to a hosted service.

The context numbers do not come from the endpoint

This is the failure that costs the most time to diagnose, because nothing errors.

In local mode, Open Interpreter sets context_window to 3000 and max_tokens to 1000. Those are numbers suited to a small local model, and they are not inferred when you supply --api_base — so pointing at a remote gateway that serves a 200k-context model leaves you with a 3000-token window unless you set it:

interpreter --api_base "https://gw.example.com/v1" \
            --api_key "fake_key" \
            --model openai/gpt-6-sol \
            --context_window 128000 \
            --max_tokens 8192

The constraint to respect is max_tokens smaller than context_window. What you see when you get it wrong is not an error but truncation: long sessions lose their beginning, and long code output gets cut off mid-function. The model then "fixes" a truncated file it never saw the end of.

If responses are slow or fail, the documented move is to try a much shorter window first, around 1000, because smaller windows use less memory — then raise it once you know the endpoint is healthy.

Profiles instead of repeating flags

Rather than typing the same four flags every session, Open Interpreter reads YAML profiles:

interpreter --profiles        # opens the profiles directory
interpreter --profile my_gateway.yaml

default.yaml is what loads when you do not name one, so putting your endpoint, model and context numbers there makes every session start connected. This is also the only clean way to keep two endpoints — a local one for experiments and a remote one for real work — without remembering which flags go with which.

Across 219 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (150 models) charge at least 4× more for output than input. The other 624 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Why the input side grows faster here

An interpreted session has a shape no chat client has: the model writes code, the shell runs it, and the output of that execution is fed back as input on the next turn. A command that prints two thousand lines of stack trace is billed as input. The loop repeats until the task is done.

A day of work on this page's workload is 260k input and 35k output tokens — 5.2M input and 0.7M output across 20 working days. The code the model writes is small; the terminal output it has to read back is not. See output vs input pricing for how to tell which side your own workload actually sits on.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
gpt-6-astra$5.00 / $25.00$2.18$43.50
deepseek-v4-flash$0.22 / $0.66$0.08$1.61
glm-5.3$0.70 / $2.2$0.26$5.18
MiniMax-M3$0.15 / $0.60$0.06$1.20
qwen3.8-max$1.00 / $3.00$0.36$7.30

One day = 260k input + 35k output on this page's workload. 20 days = 5.2M input and 0.7M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Verifying the route

  1. Send --api_key even for a keyless server. fake_key is the documented placeholder; an omitted key fails before anything reaches your gateway.
  2. Prefix the model with openai/ in Python. Without the prefix LiteLLM guesses the provider, and the guess changes the request format.
  3. Set --context_window and --max_tokens explicitly. Nothing about --api_base tells Open Interpreter how large your model's window is, and the defaults are 3000 and 1000.
  4. Run --verbose once. It prints the request details, which is the fastest way to confirm the host, path and model id that actually went out.
  5. Ask for a long output on purpose — a directory listing of a large folder. If it comes back truncated, the context numbers are still at their defaults, not the endpoint.

FAQ

FAQ

Why is api_key required when my local server does not check it?

LiteLLM, which Open Interpreter uses to reach models, requires a non-empty key. Use fake_key, or your real key against a gateway that checks.

Do I need the openai/ prefix on the model name?

In Python, yes. LiteLLM infers the provider from the model string, and the prefix is what selects the OpenAI request format your gateway expects.

My responses are truncated but there is no error

context_window and max_tokens are still at their defaults — 3000 and 1000 in local mode. They are not inferred from --api_base. Set both, and keep max_tokens below context_window.

What does interpreter.offline do?

It disables the online features such as the hosted procedure library, so the session does not make calls beyond the endpoint you configured. It is independent of api_base.

Can I keep more than one endpoint?

Yes — save each as a YAML profile with interpreter --profiles and select with --profile. default.yaml is used when you do not name one.

How is this different from LiteLLM configuration?

Open Interpreter uses LiteLLM internally, so the provider prefixes and the key requirement come from there, but you set them on interpreter.llm rather than in a LiteLLM config. For LiteLLM's own surface see LiteLLM with a custom endpoint, and for a terminal agent with a review sandbox instead of live execution, Aider.

Related

Get API access