AutoGen with a custom endpoint
AutoGen 0.4 models an application as agents in conversation, and each agent carries its own model client. Pointing one at a different endpoint takes two arguments — but the second one, the capability declaration, is what decides whether your agents can use tools at all.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
What you need before you start
- A gateway API key and base URL (
https://aicomp.ai/v1) - Python 3.10 or later — on 3.9 pip silently installs the incompatible 0.2 release
pip install autogen-agentchat "autogen-ext[openai]"
The client, with capabilities declared
import asyncio, os
from autogen_ext.models.openai import OpenAIChatCompletionClient
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.conditions import MaxMessageTermination
client = OpenAIChatCompletionClient(
model="gpt-5.6-luna",
base_url="https://aicomp.ai/v1",
api_key=os.environ["GATEWAY_API_KEY"],
model_info={
"vision": False,
"function_calling": True, # REQUIRED for tool-using agents
"json_output": True,
"structured_output": True,
"family": "unknown",
},
)
async def main():
writer = AssistantAgent("writer", client, system_message="Draft.")
critic = AssistantAgent("critic", client, system_message="Critique.")
team = RoundRobinGroupChat([writer, critic],
termination_condition=MaxMessageTermination(6))
await team.run(task="Explain what a rate limit is, in two sentences.")
asyncio.run(main())
The model_info mapping is not decoration. AutoGen cannot probe an unknown endpoint for what the model supports, so anything you leave out is treated as unsupported. function_calling is the one that matters most: AssistantAgent and most other AgentChat agents require it, and the framework only tells you so after you attach a tool.
That ordering is what makes this hard to debug. A plain conversation succeeds — which is what everyone tests first — and then the first tool-using agent either raises or, worse, answers in prose while the workflow reports completion.
Cost: the product, not the sum
The arithmetic that catches people out is that a group chat bills one call per agent per turn. A team of three running eight turns is twenty-four calls for one task, and each turn carries the conversation history forward, so per-call input tokens grow as the run proceeds.
# A group chat does not cost one call per turn -- it costs one call per
# agent per turn. Two agents over six turns is twelve billed calls, and each
# turn carries the accumulated conversation history forward.
def autogen_calls(n_agents, n_turns):
"""Upper bound on billed calls for a RoundRobinGroupChat run."""
return n_agents * n_turns
def autogen_month_cost(n_agents, n_turns, runs_per_day,
in_tok_per_call, out_tok_per_call,
in_rate, out_rate):
calls = autogen_calls(n_agents, n_turns) * runs_per_day * 30
return calls * (
in_tok_per_call / 1_000_000 * in_rate
+ out_tok_per_call / 1_000_000 * out_rate
)
# 3 agents, 8 turns, 50 runs/day, 4k input and 500 output per call:
# 3 * 8 * 50 * 30 = 36,000 calls a month, not 1,500.
Run that before you deploy a team rather than after. The two levers it exposes are both structural rather than pricing-related: fewer agents, and a termination condition that actually terminates. A team that runs to max_turns because the termination string never matched costs the full budget every single time.
Caching is off by default in v0.4
v0.2 cached by default through cache_seed. v0.4 does not, so re-running the same conversation bills again — invisible during prompt iteration, and visible only as an unexplained line in the monthly total.
from autogen_ext.models.cache import ChatCompletionCache
from autogen_ext.cache_store.diskcache import DiskCacheStore
from diskcache import Cache
# Caching is OFF by default in v0.4 (it was on by default in v0.2).
# A re-run of the same conversation is billed again unless you opt in.
cache_client = ChatCompletionCache(
client, DiskCacheStore[CHAT_CACHE_VALUE_TYPE](Cache("/tmp/autogen-cache"))
)
What it costs
| Model | Gateway rate in / out per 1M tokens | Official list in / out per 1M tokens | Diff |
|---|---|---|---|
| gpt-5.6-luna | $0.1 / $0.6 | $0.2 / $1.2 | 50% |
| claude-sonnet-5 | $1 / $5 | $2 / $10 | 50% |
| deepseek-v4-flash | $0.22 / $0.66 | — / — | — |
Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.
Rates checked 2026-09-26. Multiply these by agents × turns × runs before comparing — a per-million-token difference of a few dollars becomes the dominant term at thirty-six thousand calls a month.
AutoGen failures and what each one is actually telling you
| Symptom | What is actually happening | How to confirm | Fix |
|---|---|---|---|
| ValueError: Model does not support function calling | function_calling was not declared in model_info, so the capability is assumed absent. | It appears only after a tool is attached — plain chat succeeds. | Declare function_calling: True and re-test with one harmless tool. |
| The agent answers in prose and the workflow reports success | Same root cause, but the agent falls back to text instead of raising. | Inspect the returned message for a tool_call attribute rather than reading the text. | Declare capabilities, and assert on the tool call object, not on the prose. |
| ModuleNotFoundError: No module named 'autogen_agentchat' | The 0.2 release was installed, usually because pip resolved it on Python 3.9. | python --version, then pip show autogen-agentchat. | Use Python 3.10+ and install autogen-agentchat with a bounded version. |
| TypeError on an unexpected keyword argument | model_info vs model_capabilities — both names appear in AutoGen's own materials across versions. | The error names the accepted keyword; read it rather than guessing. | Use whichever the installed signature accepts. |
| The run never terminates and costs the full budget | The termination condition never matched, so the team ran to max_turns. | Log the number of turns actually taken, not just the final answer. | Set max_turns as a hard ceiling and alert when it is reached. |
| The same conversation is billed twice | Caching is off by default in v0.4. | Re-run an identical input and compare the usage log. | Wrap the client in ChatCompletionCache with a disk or Redis store. |
When this is the wrong move
Cases where a multi-agent team is the wrong shape
| Situation | Why it breaks | Do this instead |
|---|---|---|
| Your agents need capabilities you cannot confirm | Undeclared capabilities are treated as absent, and the framework only reports it at tool-attach time. | Confirm tool calling, JSON output and structured output against the endpoint first. |
| The conversation history is long and every agent re-reads it | Per-turn input tokens grow across a run, so the marginal turn is the most expensive one. | Trim history between turns, or use fewer agents before optimizing model price. |
| You cannot bound the number of turns | An unbounded team is an unbounded bill; a termination string that never matches costs the ceiling every run. | Set max_turns and alert when it is reached. |
Rolling it out without finding out the hard way
The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:
- Prove it on one two-agent team with one harmless tool and max_turns set. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
- Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
- Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
- Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.
Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.
FAQ
Why does my agent reply in prose instead of calling the tool?
AutoGen cannot infer what an unknown model can do, so undeclared capabilities are treated as absent. The consequence is specifically nasty: the failure only appears once you attach a tool. A plain conversation works perfectly, which is what most people test, and then the first tool-using agent raises ValueError: Model does not support function calling — or, worse, quietly answers in prose and the workflow appears to complete. Declare function_calling, json_output, vision and structured_output in model_info, and verify by attaching one harmless tool with known input and output.
Is the parameter called model_info or model_capabilities?
Both names appear in AutoGen's own materials — the API reference lists model_info, while several examples and integration guides write model_capabilities. Rather than guess, construct the client with one and read the TypeError if it is wrong: the signature names the accepted keyword. The content is the same either way — a mapping of capability names to booleans, required whenever the model name is not a recognised OpenAI model.
Why does installing AutoGen not give me autogen_agentchat?
You have the 0.2 release. On Python 3.9, pip resolves to it silently because v0.4 and later require Python 3.10, and the 0.2 package ships a module named autogen instead of autogen_agentchat. The two APIs are not interchangeable. Check python --version before anything else — this is a five-second check that explains a lot of confusing import errors.
How do I estimate what a group chat will cost?
Multiply. A RoundRobinGroupChat bills one call per agent per turn, and every turn carries the accumulated history forward, so input tokens grow across the conversation. Two agents over six turns is twelve calls, not six; three agents over eight turns run fifty times a day is thirty-six thousand calls a month. The snippet above is the arithmetic worth running before you deploy a team rather than after.
Is response caching on by default?
No — v0.4 turned it off. v0.2 enabled it by default via cache_seed; in v0.4 you opt in by wrapping the client in ChatCompletionCache with a disk or Redis store. The practical effect is that re-running the same conversation bills again, which is easy to miss while iterating on prompts and only shows up as an unexplained line in the monthly total.