Spring AI with a Custom OpenAI-Compatible Endpoint

In short: spring.ai.openai.base-url is a prefix with no version segment — the documented default is https://api.openai.com and completions-path supplies /v1/chat/completions, so pasting a /v1 URL into base-url produces a doubled path. api-key is mandatory even against a server that does not authenticate, maxTokens and maxCompletionTokens cannot both be set, and the starter artifact was renamed.

Spring AI's OpenAI starter is the shortest path from a Spring Boot service to any OpenAI-shaped gateway. Add one dependency, set two properties, and the auto-configuration hands you a ChatClient.Builder. Almost every part of that sentence has a trap in it.

spring.ai.openai.base-url does not contain /v1. The version prefix lives in spring.ai.openai.chat.completions-path, whose documented default is /v1/chat/completions and whose job is to be "appended to the base URL". Paste a /v1-suffixed URL into base-url, as several tutorials instruct, and the request goes to /v1/v1/chat/completions. api-key is required even when your gateway does not authenticate — and a placeholder left blank is still sent as a bearer token. And the starter artifact was renamed, so a pom copied from a 2024 tutorial will not resolve at all.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page covers the two properties that decide routing, why the version prefix belongs in the path rather than the URL, what each of the three ways to override them actually scopes, and what a server-side application costs per month once retrieved context is counted.

The four properties that matter

PropertyDefaultWhat it controls
spring.ai.openai.base-urlhttps://api.openai.comscheme, host and port only
spring.ai.openai.chat.completions-path/v1/chat/completionsthe path appended to the base URL
spring.ai.openai.api-key(none)the bearer token, sent unconditionally
spring.ai.openai.chat.options.modelgpt-4o-minithe model field, sent verbatim

Auto-configuration builds an OpenAiChatModel and a ChatClient.Builder as soon as an API key is present. No @Bean is needed for the basic case.

spring:
  ai:
    openai:
      base-url: https://gw.example.com
      api-key: ${GATEWAY_API_KEY}
      chat:
        options:
          model: claude-sonnet-5

base-url is not a URL you paste in

The default value proves the shape: https://api.openai.com. No /v1, no path segment at all — because the version is already in the completions path that gets concatenated onto it.

So the field is a prefix, and the destination is base-url + completions-path. That has three consequences:

api-key is mandatory, including when it is not used

Spring AI's OpenAI client always sends an authorization header. A gateway that expects no credential has no way to signal that, so the documented approach is a placeholder:

spring.ai.openai.api-key: "not-used"

Two failure modes around this field are worth naming:

The starter artifact was renamed

The dependency coordinates changed. Older guides reference:

<artifactId>spring-ai-openai-spring-boot-starter</artifactId>

The current starter is:

<dependency>
  <groupId>org.springframework.ai</groupId>
  <artifactId>spring-ai-starter-model-openai</artifactId>
</dependency>

A copied pom against a recent BOM fails at resolution, before any of the routing questions arise. Spring AI also publishes milestones, so an older 1.0.0-M* coordinate will pull an old property layout with it — check the BOM version you actually resolved, not the one in the tutorial.

Two options that cannot coexist

For output length, Spring AI exposes both maxTokens and maxCompletionTokens, and the documentation is unusually direct: they are mutually exclusive, and setting both results in an API error. maxTokens is for non-reasoning models, maxCompletionTokens for reasoning models. Set the one that matches the model and remove the other — a YAML file that sets both because both looked plausible produces a 400 that reads like a gateway problem.

Across 219 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (150 models) charge at least 4× more for output than input. The other 624 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Why the input column dominates here

A Spring AI service is usually a retrieval service: documents, chunks or tool results are assembled into the prompt, and the answer is a few paragraphs. Input scales with the size of the retrieved context, not with the length of the question, and every request starts from zero — nothing an editor or a chat client would carry in a session.

A day of work on this page's workload is 280k input and 45k output tokens — 5.6M input and 0.9M output across 20 working days. With output billed at a multiple of input on most gateways, the ratio is what makes or breaks the monthly number, which is why the input rate is the one to negotiate.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-sonnet-5$1.00 / $5.00$0.51$10.10
gpt-5.6-sol$2.5 / $15.00$1.38$27.50
deepseek-v4-pro$0.66 / $1.98$0.27$5.48
kimi-k3$1.5 / $7.5$0.76$15.15
qwen3.8-max$1.00 / $3.00$0.42$8.30

One day = 280k input + 45k output on this page's workload. 20 days = 5.6M input and 0.9M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Retries are on by default

Spring AI configures a retry policy under spring.ai.retry, and the documented default is generous — spring.ai.retry.max-attempts is 10. Against a gateway that returns 429 under load, that is ten attempts per call with exponential backoff. Set it deliberately rather than inheriting it, and set spring.ai.retry.on-client-errors with intent: the default does not retry 4xx responses, so a wrong model ID fails immediately while a slow gateway gets retried ten times. That asymmetry is the correct default, and it is worth knowing before you debug a request that seems to hang.

Spring AI also sends a User-Agent: spring-ai header on every request. That is the fastest way to confirm, from the gateway side, that the traffic you are looking at came from this application and not from another client pointed at the same key.

Verifying the route

  1. Start the application and watch for the auto-configuration. A missing or empty api-key is reported at startup, which is a far cheaper failure than a 401 later.
  2. curl the base URL with the completions path appended, by hand: ${BASE_URL}/v1/chat/completions. A valid-shaped JSON error proves the route exists; an HTML page proves it does not.
  3. Send one message through ChatClient and check the gateway log for the path. One /v1 is correct, two means base-url carried the prefix, zero means completions-path was emptied.
  4. Ask for a long completion with a low maxTokens, then the same with maxCompletionTokens on a reasoning model. That is the pair that proves which of the two exclusive options your model needs.
  5. Confirm the model field. spring.ai.openai.chat.options.model is transmitted verbatim, so a 404 from an otherwise healthy route points at the ID rather than the URL.

FAQ

FAQ

Should `base-url` include `/v1`?

No. The documented default is https://api.openai.com with no path, and completions-path supplies /v1/chat/completions. Adding the prefix to both produces /v1/v1/chat/completions.

Can I run Spring AI against a server with no authentication?

Yes, with a placeholder API key. The property is mandatory, so set it to a string such as "not-used" rather than leaving it out — and make sure no stale value is still being injected from the environment.

Which property scopes my gateway to chat only?

spring.ai.openai.chat.base-url. The outer spring.ai.openai.base-url is connection-level and is shared by every OpenAI-backed client the auto-configuration creates, including embeddings.

Why does my pom fail to resolve the starter?

The artifact was renamed. spring-ai-openai-spring-boot-starter is the old coordinate; current applications use spring-ai-starter-model-openai, and the property layout differs between the two.

Why do I get an API error when I set an output limit?

maxTokens and maxCompletionTokens are mutually exclusive and setting both is documented to produce an error. Use maxTokens for non-reasoning models and maxCompletionTokens for reasoning models.

Is Spring AI cheaper than using an HTTP client directly?

The price is the endpoint's, not the library's — Spring AI is a client. What changes the bill is context assembly: a retrieval service re-sends retrieved documents on every request, which is why this page's workload is input-heavy. For the same pattern in a different framework, see Semantic Kernel on the .NET and Python side.

Related

Get API access