Agno with a Custom OpenAI-Compatible Endpoint

In short: OpenAILike takes base_url with the version segment and an explicit id, because both id and api_key default to the literal string not-provided and the first error arrives from the gateway rather than from your code. The knowledge base is a second client: its embedder defaults to OpenAI's own endpoint, so an agent entirely on your gateway can still be sending every document elsewhere.

Agno is a Python framework for multi-agent systems with memory, knowledge and tools. It is built around a small API surface and very fast agent start-up, and every OpenAI-shaped gateway connects through one class: OpenAILike. Two lines get you connected, and two defaults are what make those two lines fail quietly.

Both id and api_key default to the literal string not-provided. Nothing raises at construction time — the request goes out carrying a model name that no gateway recognises, and the error arrives from the server rather than from your own code. And OpenAILike covers the chat model only: the moment you attach a knowledge base, its embedder defaults to OpenAI's embedding service, so an agent that is entirely on your gateway can still be sending every document to a second vendor you did not configure.

You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment

This page covers the import that actually works, the two defaults, the separate embedder, and what a month of multi-agent traffic costs once every agent is making its own calls.

The connection

Install both packages — Agno's OpenAI-compatible models need the OpenAI SDK present in the environment:

pip install agno openai

Then build the model and hand it to an agent:

import os
from agno.agent import Agent
from agno.models.openai.like import OpenAILike

agent = Agent(
    model=OpenAILike(
        id="gpt-6-sol",
        base_url="https://gw.example.com/v1",
        api_key=os.environ["MY_GATEWAY_API_KEY"],
    ),
    markdown=True,
)
agent.print_response("Explain what an LLM gateway does in one sentence.")

base_url is the endpoint root and includes the version segment. id is the model identifier your gateway accepts, which is not necessarily the display name — gateways that serve several vendors usually expect a provider-prefixed slug, so read the ID off the gateway's model list rather than off a pricing page.

OpenAILike accepts everything OpenAIChat accepts, so temperature, structured output and tool calling all behave the same way once the endpoint is right.

Two defaults named not-provided

The parameter table lists id with a default of not-provided and api_key with the same. That is a real string, not a sentinel that triggers validation:

OpenAILike(base_url="https://gw.example.com/v1")   # no error here

This constructs fine. The failure happens on the first request, and it comes back as a model or auth error from the gateway — which reads like a gateway problem. Passing both explicitly costs nothing and moves the error to where you can see it.

The second half of the same trap is the SDK. Agno's OpenAI-compatible classes route through the OpenAI SDK, so a missing openai package surfaces as ModuleNotFoundError: No module named 'openai' on import, even though you installed Agno.

The embedder is a separate configuration

This is the one that costs money rather than time. An agent with a knowledge base has two model paths, and only one of them is the object you configured:

from agno.agent import Agent
from agno.knowledge import PDFUrlKnowledgeBase
from agno.models.openai.like import OpenAILike
from agno.embedder.openai import OpenAIEmbedder
from agno.vectordb.pgvector import PgVector

knowledge = PDFUrlKnowledgeBase(
    urls=["https://example.com/spec.pdf"],
    vector_db=PgVector(
        table_name="docs",
        db_url=db_url,
        embedder=OpenAIEmbedder(                     # ← separate endpoint
            api_key=os.environ["MY_GATEWAY_API_KEY"],
            base_url="https://gw.example.com/v1",
        ),
    ),
)

agent = Agent(model=my_model, knowledge=knowledge)

Leave the embedder out and it defaults to OpenAIEmbedder pointed at OpenAI's own endpoint. The symptom is an authentication error mentioning OpenAI while your chat calls are working perfectly — which reads as a contradiction until you know there are two clients. It also means every document you ingest is billed by a vendor you did not intend to use.

The rule that follows: for each component you add — chat model, embedder, and anything else that makes requests — check whether it inherited a default endpoint, and set it explicitly.

Across 219 catalogue models that publish both rates, the median output rate is 4.0× the input rate, and 68% of them (150 models) charge at least 4× more for output than input. The other 624 raw catalogue entries are dropped by cleaning — test entries, dated snapshots and non-chat variants — not for a missing rate. This is why an agent workload is decided on the output column, not the input one everyone quotes.

Many agents means many calls

Agno's design point is cheap agent instantiation, and that makes it easy to build a system where a dozen agents each make a handful of calls per task. The per-call cost is small; the count is not. Two things dominate:

A day of work on this page's workload is 320k input and 60k output tokens — 6.4M input and 1.2M output across 20 working days. For the framework-level view of where that money goes, see how to monitor API spend, and for the output side of the same arithmetic, output vs input pricing.

Per 1M input / output tokens and a 20-day month
ModelGateway rate
in / out per 1M tokens
One day20 days
claude-opus-5$2.5 / $12.5$1.55$31.00
gpt-5.6-terra$1.00 / $6.00$0.68$13.60
kimi-k3$1.5 / $7.5$0.93$18.60
deepseek-v4-pro$0.66 / $1.98$0.33$6.60
gemini-3.7-flash$0.375 / $1.875$0.23$4.65

One day = 320k input + 60k output on this page's workload. 20 days = 6.4M input and 1.2M output. Rates checked 2026-10-10 — gateway rates move with upstream promotions, so verify the current number in your dashboard before committing to a budget.

Verifying the route

  1. Print the model id before the first call. If it says not-provided, the request will fail at the gateway, not in your code.
  2. Ask the gateway for its model list and copy the id from there. Display names and gateway ids differ, especially on multi-vendor gateways.
  3. Watch the gateway log for a single request. One entry means the chat model is routed; a second host appearing when you attach a knowledge base means the embedder is still on its default.
  4. Ingest one small document and confirm the embedding request also lands on your gateway. This is the check that catches the split configuration.
  5. Run with a trivial prompt first — one sentence, no tools, no knowledge. It isolates the endpoint from everything else.

FAQ

FAQ

Why does it fail with a model error when I never set a model?

id defaults to the string not-provided. There is no validation at construction, so the gateway is the first thing to reject it. Always pass id explicitly.

Do I need the OpenAI SDK if I am not using OpenAI?

Yes. Agno's OpenAI-compatible classes are built on it. Without the package you get ModuleNotFoundError: No module named 'openai' at import.

Should base_url include /v1?

Yes. It is the endpoint root including the version segment — https://gw.example.com/v1. Agno does not append the version path for you.

Why does adding a knowledge base produce OpenAI auth errors?

The embedder defaults to OpenAI's endpoint independently of your chat model. Configure the embedder with your own base_url and api_key; OpenAILike on the agent does not cover it.

Can several agents share one gateway key?

Yes. The key controls access and budget, and each OpenAILike instance picks its own model, so different agents can use different models behind one credential.

How is this different from CrewAI or AutoGen?

The endpoint mechanics are the same — a base URL and a model id per model object. The difference is the failure surface: Agno's defaults are silent strings rather than exceptions, and its knowledge components carry their own endpoint. For the crew-shaped version of the same setup see CrewAI, and for the SDK-level view, OpenAI & Anthropic SDK with a custom endpoint.

Related

Get API access