Gemini CLI with a custom endpoint

Google's terminal agent can run against any OpenAI-compatible endpoint. Two fields in settings.json and one auth selection do it — but the failure mode that costs people an afternoon is not a 401. It is a CLI that answers you in perfect prose and never touches a single file.

In short: Gemini CLI performs every file edit and shell command through function calling, so a model without tool-call support does not error — it replies in text and never modifies the repo. Test with a reversible one-line change.
You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment
Confirm the key works first. One command, no SDK, costs nothing:
curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.

What you need before you start

Step 1 — the api block in settings.json

The file is ~/.gemini/settings.json for user-wide settings, or .gemini/settings.json in a project if you want the setting to travel with the repository.

{
  "api": {
    "baseUrl": "https://aicomp.ai/v1",
    "apiKey": "sk-your-gateway-key"
  }
}

Step 2 — run /auth and choose the right surface

This is the step people skip. Without an explicit auth selection the CLI assumes Google OAuth and prompts you to log in — it presents as a login request, not as a configuration error, so it is easy to conclude the endpoint is wrong when nothing was ever sent to it. Start the CLI, run /auth, and select the OpenAI-compatible option.

Step 3 — pick a model that can call tools

Everything this CLI does that is useful — editing files, running shell commands, reading the repo — is a function call. A model without tool-call support does not fail; it simply answers conversationally, and you get a session that looks like it is working and changes nothing on disk.

Test it with something small and reversible. Ask for a one-line change to a file you can restore. If the answer is correct but the file is untouched, the tool path is not working, and no amount of base-URL tweaking will fix it.

The environment-variable route

For CI, or for a quick test, the variables are quicker than editing the file:

export GEMINI_API_KEY="sk-your-gateway-key"
export GOOGLE_GEMINI_BASE_URL="https://aicomp.ai/v1"

gemini -p "summarise this repository"

Be deliberate about which surface you are using. These variables belong to the native Gemini protocol path; the settings.json api block belongs to the OpenAI-compatible path. Configuring both and expecting one behaviour is how you end up authenticated against a surface you did not intend.

What it costs

Models reachable from Gemini CLI — gateway rate vs official list
ModelGateway rate
in / out per 1M tokens
Official list
in / out per 1M tokens
Diff
gemini-3.7-flash$0.375 / $1.875— / ——
gpt-5.6-luna$0.1 / $0.6$0.2 / $1.250%
deepseek-v4-flash$0.22 / $0.66— / ——

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.

The cost shape here is distinctive: the first turn of a session carries the repository context and is by far the most expensive request of the run; subsequent turns are cheap by comparison. That means session count matters more than turn count. Ten short sessions on the same repo cost noticeably more than one long one, even at the same total number of prompts.

Gemini CLI failures, ordered by how much time they waste

SymptomWhat is actually happeningHow to confirmFix
Keeps prompting for Google loginAuth type not selected; the CLI defaults to OAuth and never calls your endpoint.Check whether any request reached the usage log — usually none has.Run /auth and pick the OpenAI-compatible option after setting the api block.
Answers but never edits filesThe model does not support function calling, and the CLI does not report that as an error.Ask for a one-line change to a file you can restore; if nothing changes, this is it.Switch to a model that exposes tool calls.
Works in one terminal, not anotherEnvironment variables are per-shell; the CLI reads what its own process inherited.Print the variables in the same shell that launches the CLI.Put the settings in settings.json instead of relying on the shell.
401 with a key that works elsewhereWrong surface — the key is being validated against a different API than the one you configured.Compare the error body against a raw curl call to the same path.Configure one surface only: either the api block or the native variables, not both.
404 or a path that looks doubledThe base URL version segment does not match what the surface expects.Read the path in the error output — the duplication or omission is visible in it.Match the base URL to the surface you selected in step 2.

When this is the wrong move

When this setup is not worth it

SituationWhy it breaksDo this instead
Your model cannot call toolsThe CLI's entire value — edits, shell, repo reads — is function calling.Use a chat-oriented tool instead, or pick a model that exposes tool calls.
You want the setting to follow the repoUser-wide settings do not travel with a checkout.Use .gemini/settings.json in the project instead of the home-directory file.
You are comparing providers on price aloneSession count, not prompt count, drives this bill — a rate table will mislead you.Measure tokens per session on your own repo before comparing.

Rolling it out without finding out the hard way

The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:

  1. Prove it on one reversible one-line file edit. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
  2. Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
  3. Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
  4. Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.

Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.

FAQ

Why does Gemini CLI keep asking me to log in with Google?

Because the auth type was never set. The CLI falls back to Google OAuth whenever it does not find the credentials and endpoint it expects, which presents as a login prompt rather than as an error. Set api.baseUrl and api.apiKey, then run /auth inside the CLI and select the OpenAI-compatible option. If the variables are set but the prompt still appears, they were exported in a different shell than the one that launched the CLI.

The CLI answers my questions but never edits files. Why?

Every file edit and every shell command goes through function calling. A model that does not expose tool calls will still answer conversationally — you get plausible text and no changes, with no error anywhere. This is the most common and the least obvious failure mode on this tool. Switch to a model that supports tool calls and re-test with a change small enough to revert.

Should I use settings.json or environment variables?

The file is the reliable route for the OpenAI-compatible path, because it is what the auth selection reads. Environment variables — GEMINI_API_KEY and GOOGLE_GEMINI_BASE_URL — are quicker for a one-off and are what you want in CI, but they belong to the native Gemini protocol path rather than the OpenAI-compatible one. Use one or the other deliberately; mixing them is how you end up authenticated against the wrong surface.

Does the base URL need a /v1 suffix?

It depends on which surface you are pointing at, and this is worth checking rather than guessing. The OpenAI-compatible path expects a versioned base; the native path does not. The CLI will report the resulting path in its error output, and a doubled or missing segment is visible there.

Why is my first run slow but later ones fast?

The first call in a session carries the repository context, which is large. Later turns in the same session reuse it. That shape — a heavy first request followed by cheaper ones — is also why the per-token input rate matters more than the output rate on this tool.

Related

Get API access