CrewAI with an OpenAI-compatible endpoint
CrewAI runs teams of agents, and each one is its own caller. That makes the cost shape multiplicative rather than additive — agents times turns — which is the thing a per-token rate table cannot show you.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
Step 1 — build the LLM object
from crewai import LLM, Agent
llm = LLM(
model="openai/gpt-5.6-luna",
base_url="https://aicomp.ai/v1",
api_key="sk-your-gateway-key",
temperature=0.2,
)
researcher = Agent(
role="Researcher",
goal="Answer precisely, with sources",
backstory="You check before you answer.",
llm=llm,
)
The object is what carries the endpoint. Handing it to an agent is what makes that agent use it; anything constructed without it falls back to the framework's defaults.
Step 2 — the flag that matters: custom_openai
CrewAI validates model identifiers against a known list. An ID served by a gateway is frequently not on that list, and the result is a rewrite or a refusal rather than a clean error. custom_openai=True tells the framework to pass the identifier through untouched.
# custom_openai=True lets a gateway model ID through unrewritten.
# Without it, CrewAI checks the ID against OpenAI's own model list.
llm = LLM(
model="openai/deepseek-v4-flash",
custom_openai=True,
base_url="https://aicomp.ai/v1",
api_key="sk-your-gateway-key",
)
The openai/ prefix selects the OpenAI-compatible route; the flag protects the ID itself. You generally want both when the model is not an OpenAI model.
Step 3 — agents defined in YAML
A base URL cannot be written as a string in agents.yaml. Build the object in a decorated method and reference it by name:
from crewai.project import CrewBase, llm
@CrewBase
class ResearchCrew:
@llm
def gateway_llm(self):
return LLM(
model="openai/gpt-5.6-luna",
custom_openai=True,
base_url="https://aicomp.ai/v1",
api_key="sk-your-gateway-key",
)
# agents.yaml then references it by name:
# researcher:
# llm: gateway_llm
This is also the tidier place for credentials — the YAML holds a name, not a key.
What a crew actually costs
| Crew shape | gpt-5.6-luna | claude-sonnet-5 | deepseek-v4-flash |
|---|---|---|---|
| 2 × 4 Two agents, four turns each — 8 calls | $0.00976 | $0.092 | $0.017776 |
| 4 × 6 Four agents, six turns each — 24 calls | $0.02928 | $0.276 | $0.053328 |
| 6 × 10 Six agents, ten turns each — 60 calls | $0.0732 | $0.69 | $0.13332 |
Agents multiply, and so do turns — a crew of six running ten turns each is sixty calls, not one. This table is the part of CrewAI cost that a per-token rate cannot tell you.
The multiplication is the whole story. A single call priced at $0.00122 is trivial; the same call made sixty times inside a six-agent, ten-turn crew is not. Before switching models, look at whether the crew shape itself can be smaller — fewer agents, fewer turns, or a cheaper model for the agents doing mechanical work.
CrewAI failures, and the one that costs real money
| Symptom | What is actually happening | How to confirm | Fix |
|---|---|---|---|
| Model ID rejected or silently rewritten | CrewAI validates identifiers against a known provider list; a gateway ID is not on it. | Try the same ID with a raw curl call — if that works and CrewAI does not, validation is the cause. | Set custom_openai=True and prefix the ID with openai/. |
| Agents from YAML ignore the endpoint | A plain string in agents.yaml cannot carry a base URL. | Check whether the agent's calls appear in the usage log at all. | Define the LLM in an @llm-decorated method and reference it by name. |
| Cost far above a single-call estimate | Agents multiply by turns; each call carries its own context. | Count entries in the usage log for one crew run and compare with your agent count. | Reduce agents or turns before switching models. |
| One agent works, another does not | Each agent is configured separately; those without an llm argument use defaults. | Compare which model IDs appear in the usage log. | Pass the LLM object explicitly to every agent. |
| Environment variables appear to be ignored | Explicit arguments to LLM() take precedence over the environment. | Print the resolved configuration, or remove the argument and retest. | Use one mechanism consistently — arguments in code, variables in containers. |
When this is the wrong move
Where crews cost more than they return
| Situation | Why it breaks | Do this instead |
|---|---|---|
| You are adding agents to improve output | Each agent multiplies the call count; the gain is rarely proportional. | Add an agent only when a task genuinely needs a separate role. |
| Your tasks are single-step | Crew overhead — multiple agents, multiple turns — is wasted on one-shot work. | Call the model directly. |
| You have not counted calls per run yet | Without that number, every cost change is guesswork. | Count usage-log entries for one run, then optimise. |
How to confirm the change actually took effect
"I set it and nothing happened" is the most common failure here, and it is almost never the endpoint. It is usually the client: a desktop app that reads its own settings file rather than your shell environment, a process started before you exported the variable, or an SDK client instance built once at import time and reused after the config changed. Verify before you debug anything else.
Layer 1 — is the endpoint reachable, and does your key work?
curl -s https://aicomp.ai/v1/models \
-H "Authorization: Bearer $KEY" | head -c 400
A JSON list of model IDs means both are true. An Invalid token body means the
endpoint answered but rejected the credential — the URL is right, the key is not. An HTML page
means something between you and the endpoint (a corporate proxy, a captive portal) answered
instead, and no amount of config change will fix it from inside your app.
Layer 2 — who actually answered?
from openai import OpenAI
client = OpenAI(base_url="https://aicomp.ai/v1", api_key=KEY)
r = client.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "Reply with the single word: ok"}],
max_tokens=8,
)
print(r.model) # 回显的 model 是谁 —— 这是端点身份的唯一硬证据
print(r.usage) # prompt_tokens / completion_tokens
The model field echoed back is the only hard evidence of identity. A request that
silently fell through to a default provider still returns HTTP 200 and still returns sensible
text, which is exactly why this layer exists. If the echoed model is not the one you asked for,
you are talking to someone else.
Layer 3 — did it land in the usage log?
Reachability is not billing. Open the provider's usage view and look for the model ID you just called, on today's date. If the call succeeded but does not appear, you either hit a cached response or you are still on the old path. This layer is also the one that catches the case where two environments share a key and you are reading the wrong one.
Three ways this check lies to you
- A 200 that is not from the API. Proxies and error pages return 200 with HTML. Check that the body parses as JSON before you trust the status code.
- A stale process. Long-running servers and editors capture configuration at startup. Restart the process, not just the shell.
- A client built before the change. SDK clients snapshot
base_urlat construction. If yours is a module-level singleton, the new value never reaches it.
FAQ
Why does my model ID get rejected or rewritten?
CrewAI checks model identifiers against a known provider list, and one it does not recognise can be rewritten or refused. Setting custom_openai=True tells it to pass the identifier through unchanged, which is what you want whenever the model is served by a gateway rather than by OpenAI directly. Prefixing the ID with openai/ routes it down the OpenAI-compatible path; the flag is what stops the ID itself being altered.
Can agents come from YAML?
Yes, but a base URL cannot be expressed as a plain string in agents.yaml. Define the LLM in a method decorated with @llm inside your @CrewBase class and reference the method name from the YAML. CrewAI resolves the name to the object when it loads the agent, which is also the cleanest place to keep credentials out of the YAML entirely.
Can different agents use different models?
Yes, and it is often the right call. Give the agent doing mechanical extraction a cheap model and the one doing synthesis a stronger one. Each agent takes its own llm argument, so a crew can mix models from the same endpoint with no extra integration work.
Why is my crew so much more expensive than a single call?
Because a crew is a product, not a request. Every agent runs every turn, so the call count is agents multiplied by turns, and each call carries its own context. A six-agent crew running ten turns is sixty calls. Reducing the number of agents usually saves more than switching to a cheaper model.
What are the environment variables?
OPENAI_API_KEY, OPENAI_API_BASE and OPENAI_MODEL_NAME. They are convenient in containers and CI, and they are read when the corresponding arguments are not passed to LLM(). Explicit arguments win where both are present, which is worth knowing before you debug a variable that is being ignored.