Spring AI with a Custom OpenAI-Compatible Endpoint
Spring AI's OpenAI starter is the shortest path from a Spring Boot service to any OpenAI-shaped gateway. Add one dependency, set two properties, and the auto-configuration hands you a ChatClient.Builder. Almost every part of that sentence has a trap in it.
spring.ai.openai.base-url does not contain /v1. The version prefix lives in spring.ai.openai.chat.completions-path, whose documented default is /v1/chat/completions and whose job is to be "appended to the base URL". Paste a /v1-suffixed URL into base-url, as several tutorials instruct, and the request goes to /v1/v1/chat/completions. api-key is required even when your gateway does not authenticate — and a placeholder left blank is still sent as a bearer token. And the starter artifact was renamed, so a pom copied from a 2024 tutorial will not resolve at all.
https://aicomp.ai/v1).
Create one free →
This page covers the two properties that decide routing, why the version prefix belongs in the path rather than the URL, what each of the three ways to override them actually scopes, and what a server-side application costs per month once retrieved context is counted.
The four properties that matter
| Property | Default | What it controls |
|---|---|---|
spring.ai.openai.base-url | https://api.openai.com | scheme, host and port only |
spring.ai.openai.chat.completions-path | /v1/chat/completions | the path appended to the base URL |
spring.ai.openai.api-key | (none) | the bearer token, sent unconditionally |
spring.ai.openai.chat.options.model | gpt-4o-mini | the model field, sent verbatim |
Auto-configuration builds an OpenAiChatModel and a ChatClient.Builder as soon as an API key is present. No @Bean is needed for the basic case.
spring:
ai:
openai:
base-url: https://gw.example.com
api-key: ${GATEWAY_API_KEY}
chat:
options:
model: claude-sonnet-5
base-url is not a URL you paste in
The default value proves the shape: https://api.openai.com. No /v1, no path segment at all — because the version is already in the completions path that gets concatenated onto it.
So the field is a prefix, and the destination is base-url + completions-path. That has three consequences:
- A URL ending in
/v1produces a doubled version segment. The request lands on/v1/v1/chat/completions. Some gateways tolerate the extra segment and some 404, which is why the mistake spreads: it works for the person who wrote the tutorial and fails for you. - A non-standard prefix is fixed in
completions-path, not inbase-url. If your gateway serves chat at/openai/v1/chat/completions, overridecompletions-pathand leavebase-urlas host and port. Jamming the prefix intobase-urlworks only by accident and breaks the moment a second path is added. base-urlis connection-level, so it is shared. Embeddings, transcription and moderation clients are configured from the same prefix.spring.ai.openai.chat.base-urloverrides it for chat only — which is what you want if the gateway serves chat completions but not embeddings, and is the reason to prefer the chat-scoped property when you are routing just one capability.
api-key is mandatory, including when it is not used
Spring AI's OpenAI client always sends an authorization header. A gateway that expects no credential has no way to signal that, so the documented approach is a placeholder:
spring.ai.openai.api-key: "not-used"
Two failure modes around this field are worth naming:
- Leaving the property out entirely fails the auto-configuration, or fails requests depending on version. The placeholder exists because the property is not optional.
- A leftover value is still a bearer token. If you switch a deployment from a real provider to an unauthenticated local server and blank the value, check the log rather than assuming — a stale key from an environment variable is sent as-is, and the error you get back is an authentication error from a server that never wanted a token.
The starter artifact was renamed
The dependency coordinates changed. Older guides reference:
<artifactId>spring-ai-openai-spring-boot-starter</artifactId>
The current starter is:
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-starter-model-openai</artifactId>
</dependency>
A copied pom against a recent BOM fails at resolution, before any of the routing questions arise. Spring AI also publishes milestones, so an older 1.0.0-M* coordinate will pull an old property layout with it — check the BOM version you actually resolved, not the one in the tutorial.
Two options that cannot coexist
For output length, Spring AI exposes both maxTokens and maxCompletionTokens, and the documentation is unusually direct: they are mutually exclusive, and setting both results in an API error. maxTokens is for non-reasoning models, maxCompletionTokens for reasoning models. Set the one that matches the model and remove the other — a YAML file that sets both because both looked plausible produces a 400 that reads like a gateway problem.
Why the input column dominates here
A Spring AI service is usually a retrieval service: documents, chunks or tool results are assembled into the prompt, and the answer is a few paragraphs. Input scales with the size of the retrieved context, not with the length of the question, and every request starts from zero — nothing an editor or a chat client would carry in a session.
A day of work on this page's workload is 280k input and 45k output tokens — 5.6M input and 0.9M output across 20 working days. With output billed at a multiple of input on most gateways, the ratio is what makes or breaks the monthly number, which is why the input rate is the one to negotiate.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-sonnet-5 | $1.00 / $5.00 | $0.51 | $10.10 |
| gpt-5.6-sol | $2.5 / $15.00 | $1.38 | $27.50 |
| deepseek-v4-pro | $0.66 / $1.98 | $0.27 | $5.48 |
| kimi-k3 | $1.5 / $7.5 | $0.76 | $15.15 |
| qwen3.8-max | $1.00 / $3.00 | $0.42 | $8.30 |
One day = 280k input + 45k output on this page's workload. 20 days = 5.6M input and 0.9M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
Retries are on by default
Spring AI configures a retry policy under spring.ai.retry, and the documented default is generous — spring.ai.retry.max-attempts is 10. Against a gateway that returns 429 under load, that is ten attempts per call with exponential backoff. Set it deliberately rather than inheriting it, and set spring.ai.retry.on-client-errors with intent: the default does not retry 4xx responses, so a wrong model ID fails immediately while a slow gateway gets retried ten times. That asymmetry is the correct default, and it is worth knowing before you debug a request that seems to hang.
Spring AI also sends a User-Agent: spring-ai header on every request. That is the fastest way to confirm, from the gateway side, that the traffic you are looking at came from this application and not from another client pointed at the same key.
Verifying the route
- Start the application and watch for the auto-configuration. A missing or empty
api-keyis reported at startup, which is a far cheaper failure than a 401 later. curlthe base URL with the completions path appended, by hand:${BASE_URL}/v1/chat/completions. A valid-shaped JSON error proves the route exists; an HTML page proves it does not.- Send one message through
ChatClientand check the gateway log for the path. One/v1is correct, two meansbase-urlcarried the prefix, zero meanscompletions-pathwas emptied. - Ask for a long completion with a low
maxTokens, then the same withmaxCompletionTokenson a reasoning model. That is the pair that proves which of the two exclusive options your model needs. - Confirm the model field.
spring.ai.openai.chat.options.modelis transmitted verbatim, so a 404 from an otherwise healthy route points at the ID rather than the URL.
FAQ
FAQ
Should `base-url` include `/v1`?
No. The documented default is https://api.openai.com with no path, and completions-path supplies /v1/chat/completions. Adding the prefix to both produces /v1/v1/chat/completions.
Can I run Spring AI against a server with no authentication?
Yes, with a placeholder API key. The property is mandatory, so set it to a string such as "not-used" rather than leaving it out — and make sure no stale value is still being injected from the environment.
Which property scopes my gateway to chat only?
spring.ai.openai.chat.base-url. The outer spring.ai.openai.base-url is connection-level and is shared by every OpenAI-backed client the auto-configuration creates, including embeddings.
Why does my pom fail to resolve the starter?
The artifact was renamed. spring-ai-openai-spring-boot-starter is the old coordinate; current applications use spring-ai-starter-model-openai, and the property layout differs between the two.
Why do I get an API error when I set an output limit?
maxTokens and maxCompletionTokens are mutually exclusive and setting both is documented to produce an error. Use maxTokens for non-reasoning models and maxCompletionTokens for reasoning models.
Is Spring AI cheaper than using an HTTP client directly?
The price is the endpoint's, not the library's — Spring AI is a client. What changes the bill is context assembly: a retrieval service re-sends retrieved documents on every request, which is why this page's workload is input-heavy. For the same pattern in a different framework, see Semantic Kernel on the .NET and Python side.
Related
- Semantic Kernel — the same server-side framework problem solved per language
- LlamaIndex and DSPy — Python retrieval stacks with the same input-heavy profile
- Tools that take one endpoint — the full matrix of clients with a custom base URL field
- Official list prices vs gateway rates — which of the two numbers on this page is which