Agno with a Custom OpenAI-Compatible Endpoint
Agno is a Python framework for multi-agent systems with memory, knowledge and tools. It is built around a small API surface and very fast agent start-up, and every OpenAI-shaped gateway connects through one class: OpenAILike. Two lines get you connected, and two defaults are what make those two lines fail quietly.
Both id and api_key default to the literal string not-provided. Nothing raises at construction time — the request goes out carrying a model name that no gateway recognises, and the error arrives from the server rather than from your own code. And OpenAILike covers the chat model only: the moment you attach a knowledge base, its embedder defaults to OpenAI's embedding service, so an agent that is entirely on your gateway can still be sending every document to a second vendor you did not configure.
https://aicomp.ai/v1).
Create one free →
This page covers the import that actually works, the two defaults, the separate embedder, and what a month of multi-agent traffic costs once every agent is making its own calls.
The connection
Install both packages — Agno's OpenAI-compatible models need the OpenAI SDK present in the environment:
pip install agno openai
Then build the model and hand it to an agent:
import os
from agno.agent import Agent
from agno.models.openai.like import OpenAILike
agent = Agent(
model=OpenAILike(
id="gpt-6-sol",
base_url="https://gw.example.com/v1",
api_key=os.environ["MY_GATEWAY_API_KEY"],
),
markdown=True,
)
agent.print_response("Explain what an LLM gateway does in one sentence.")
base_url is the endpoint root and includes the version segment. id is the model identifier your gateway accepts, which is not necessarily the display name — gateways that serve several vendors usually expect a provider-prefixed slug, so read the ID off the gateway's model list rather than off a pricing page.
OpenAILike accepts everything OpenAIChat accepts, so temperature, structured output and tool calling all behave the same way once the endpoint is right.
Two defaults named not-provided
The parameter table lists id with a default of not-provided and api_key with the same. That is a real string, not a sentinel that triggers validation:
OpenAILike(base_url="https://gw.example.com/v1") # no error here
This constructs fine. The failure happens on the first request, and it comes back as a model or auth error from the gateway — which reads like a gateway problem. Passing both explicitly costs nothing and moves the error to where you can see it.
The second half of the same trap is the SDK. Agno's OpenAI-compatible classes route through the OpenAI SDK, so a missing openai package surfaces as ModuleNotFoundError: No module named 'openai' on import, even though you installed Agno.
The embedder is a separate configuration
This is the one that costs money rather than time. An agent with a knowledge base has two model paths, and only one of them is the object you configured:
from agno.agent import Agent
from agno.knowledge import PDFUrlKnowledgeBase
from agno.models.openai.like import OpenAILike
from agno.embedder.openai import OpenAIEmbedder
from agno.vectordb.pgvector import PgVector
knowledge = PDFUrlKnowledgeBase(
urls=["https://example.com/spec.pdf"],
vector_db=PgVector(
table_name="docs",
db_url=db_url,
embedder=OpenAIEmbedder( # ← separate endpoint
api_key=os.environ["MY_GATEWAY_API_KEY"],
base_url="https://gw.example.com/v1",
),
),
)
agent = Agent(model=my_model, knowledge=knowledge)
Leave the embedder out and it defaults to OpenAIEmbedder pointed at OpenAI's own endpoint. The symptom is an authentication error mentioning OpenAI while your chat calls are working perfectly — which reads as a contradiction until you know there are two clients. It also means every document you ingest is billed by a vendor you did not intend to use.
The rule that follows: for each component you add — chat model, embedder, and anything else that makes requests — check whether it inherited a default endpoint, and set it explicitly.
Many agents means many calls
Agno's design point is cheap agent instantiation, and that makes it easy to build a system where a dozen agents each make a handful of calls per task. The per-call cost is small; the count is not. Two things dominate:
- Every agent re-sends its own context. Memory and knowledge retrieval are per agent, so a team of five working one task sends five overlapping prompts.
- Tool results come back as input. A tool that returns a large payload is billed again on the next turn, at input rates.
A day of work on this page's workload is 320k input and 60k output tokens — 6.4M input and 1.2M output across 20 working days. For the framework-level view of where that money goes, see how to monitor API spend, and for the output side of the same arithmetic, output vs input pricing.
| Model | Gateway rate in / out per 1M tokens | One day | 20 days |
|---|---|---|---|
| claude-opus-5 | $2.5 / $12.5 | $1.55 | $31.00 |
| gpt-5.6-terra | $1.00 / $6.00 | $0.68 | $13.60 |
| kimi-k3 | $1.5 / $7.5 | $0.93 | $18.60 |
| deepseek-v4-pro | $0.66 / $1.98 | $0.33 | $6.60 |
| gemini-3.7-flash | $0.375 / $1.875 | $0.23 | $4.65 |
One day = 320k input + 60k output on this page's workload. 20 days = 6.4M input and 1.2M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.
Verifying the route
- Print the model id before the first call. If it says
not-provided, the request will fail at the gateway, not in your code. - Ask the gateway for its model list and copy the id from there. Display names and gateway ids differ, especially on multi-vendor gateways.
- Watch the gateway log for a single request. One entry means the chat model is routed; a second host appearing when you attach a knowledge base means the embedder is still on its default.
- Ingest one small document and confirm the embedding request also lands on your gateway. This is the check that catches the split configuration.
- Run with a trivial prompt first — one sentence, no tools, no knowledge. It isolates the endpoint from everything else.
FAQ
FAQ
Why does it fail with a model error when I never set a model?
id defaults to the string not-provided. There is no validation at construction, so the gateway is the first thing to reject it. Always pass id explicitly.
Do I need the OpenAI SDK if I am not using OpenAI?
Yes. Agno's OpenAI-compatible classes are built on it. Without the package you get ModuleNotFoundError: No module named 'openai' at import.
Should base_url include /v1?
Yes. It is the endpoint root including the version segment — https://gw.example.com/v1. Agno does not append the version path for you.
Why does adding a knowledge base produce OpenAI auth errors?
The embedder defaults to OpenAI's endpoint independently of your chat model. Configure the embedder with your own base_url and api_key; OpenAILike on the agent does not cover it.
Can several agents share one gateway key?
Yes. The key controls access and budget, and each OpenAILike instance picks its own model, so different agents can use different models behind one credential.
How is this different from CrewAI or AutoGen?
The endpoint mechanics are the same — a base URL and a model id per model object. The difference is the failure surface: Agno's defaults are silent strings rather than exceptions, and its knowledge components carry their own endpoint. For the crew-shaped version of the same setup see CrewAI, and for the SDK-level view, OpenAI & Anthropic SDK with a custom endpoint.
Related
- CrewAI — the other multi-agent framework, configured per model object
- Pydantic AI — same idea with a different provider class
- OpenAI & Anthropic SDK — what
OpenAILikeis built on - Monitor API spend — attributing cost per run once many agents are calling
- Tools that take one endpoint — the full matrix