OpenAI Agents SDK with a custom endpoint

The Agents SDK takes a plain AsyncOpenAI client, so the base URL is one argument. Two defaults are worth knowing before you ship: tracing is on and exports to the OpenAI platform, and the SDK speaks the Responses API unless you tell it otherwise.

In short: The OpenAI Agents SDK enables tracing by default and exports it to the OpenAI platform independently of the model client, so an endpoint swap leaves traces going to api.openai.com. It also defaults to /v1/responses, not /v1/chat/completions.
You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment
Confirm the key works first. One command, no SDK, costs nothing:
curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.

What you need before you start

The working configuration

import os
from openai import AsyncOpenAI
from agents import (
    Agent, Runner, OpenAIChatCompletionsModel, set_tracing_disabled,
)

# 1. Tracing is ON by default and uploads to the OpenAI platform.
#    With no platform key in the environment, every run logs a tracing
#    failure -- your model traffic goes to the gateway, your traces do not.
set_tracing_disabled(True)

client = AsyncOpenAI(
    base_url="https://aicomp.ai/v1",
    api_key=os.environ["GATEWAY_API_KEY"],
)

# 2. The SDK defaults to the Responses API. Pick the transport explicitly.
model = OpenAIChatCompletionsModel(model="gpt-5.6-luna", openai_client=client)

agent = Agent(name="Assistant", instructions="Answer precisely.", model=model)
result = Runner.run_sync(agent, "What is a rate limit?")
print(result.final_output)

Three things are deliberate in that block. The base URL sits on the client rather than on the agent, because the agent takes a model and the model takes the client. The transport is explicit, because the default is the Responses API. And tracing is switched off, because leaving it on sends traces somewhere other than where your requests go.

Default one: tracing exports to the OpenAI platform

Tracing is enabled by default and uses the same key as your model requests unless you say otherwise. Changing the base URL moves the model traffic and leaves the exporter where it was. With no platform key present, every run logs a tracing failure — a log line that looks like a routing problem and has nothing to do with routing.

import os
from openai import AsyncOpenAI
from agents import set_default_openai_client, set_tracing_export_api_key

client = AsyncOpenAI(base_url="https://aicomp.ai/v1", api_key=os.environ["GATEWAY_API_KEY"])

# Keep the gateway key for model traffic; do NOT reuse it as the tracing key.
# Without use_for_tracing=False the SDK would send the gateway key to the
# OpenAI platform's tracing endpoint.
set_default_openai_client(client, use_for_tracing=False)
set_tracing_export_api_key(os.environ["OPENAI_PLATFORM_KEY"])

# Or keep tracing on but keep payloads out of it:
#   Runner.run(agent, "hi", run_config=RunConfig(trace_include_sensitive_data=False))
#   export OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA=0

The use_for_tracing=False argument is the one to notice. Without it the SDK may send your gateway credential to the platform's tracing endpoint, which is a credential-handling problem rather than a cost one — and it costs two lines to avoid.

Default two: the Responses API

A plain agent targets POST /v1/responses. Endpoints that serve only chat completions return 404, and a 404 on a URL you copied correctly is a genuinely confusing thing to debug.

from agents import set_default_openai_api

# Applies when the address and key come from OPENAI_BASE_URL / OPENAI_API_KEY.
set_default_openai_api("chat_completions")

# Without this, or without an explicit OpenAIChatCompletionsModel, the SDK
# targets POST /v1/responses. An endpoint that only serves chat completions
# returns 404 -- which reads like a broken base URL and is not one.

What it costs

Models reachable from the Agents SDK — gateway rate vs official list
ModelGateway rate
in / out per 1M tokens
Official list
in / out per 1M tokens
Diff
gpt-5.6-luna$0.1 / $0.6$0.2 / $1.250%
claude-haiku-4-5$0.5 / $2.5$1 / $550%
deepseek-v4-flash$0.22 / $0.66— / ——

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.

Rates checked 2026-09-26. Agent runs resend the accumulated history on every step, so input tokens dominate — the model column is worth more than any client-side change on this page.

How the bill is actually computed

Two numbers are metered separately and added: input tokens (everything you send — system prompt, conversation history, retrieved documents, tool results) and output tokens (what the model generates). Per-token, output is the expensive one — across the 216 models we track, the median output-to-input ratio is 4.0×, and the widest is 12.0×.

And yet on agent-shaped workloads the input side dominates the invoice. That is not a contradiction, it is arithmetic: an agent re-sends its entire context on every turn, so a turn that produces 800 output tokens may carry 60,000 input tokens. At a 4× unit price ratio, the input still accounts for roughly 95% of that turn's cost. Optimising output length on this workload is optimising the wrong end.

One agent turn priced at 60k input / 800 output — gateway rates, checked 2026-09-26
ModelInput / 1MOutput / 1M Output ÷ inputOne agent turnInput share of that turn
gpt-5.6-luna$0.100$0.6006.0×$0.00693%
claude-haiku-4-5-20251001$0.500$2.505.0×$0.03294%

The practical reading: unit price decides, but token shape decides harder. A model that costs twice as much per token can still be cheaper end to end if it needs half as many turns to finish the task, and a prompt that carries 40k tokens of irrelevant history costs the same as one that carries 40k tokens of useful context. Before switching models, measure which side of the sum your workload actually lives on.

Agents SDK failures and what each one is actually telling you

SymptomWhat is actually happeningHow to confirmFix
Every run logs a tracing failureTracing is on by default and exports to the OpenAI platform, configured separately from the model client.The model call succeeds; only the trace fails. That asymmetry is the tell.set_tracing_disabled(True), or set a dedicated tracing key.
404 on a base URL that works in every other toolThe SDK defaults to /v1/responses, not /v1/chat/completions.Compare the path in the request log against what the endpoint serves.Use OpenAIChatCompletionsModel or set_default_openai_api("chat_completions").
The gateway key reaches the OpenAI platformset_default_openai_client without use_for_tracing=False reuses the model key for tracing.Check which key the tracing exporter holds.Pass use_for_tracing=False and set a separate tracing key.
Customer text appears in tracesSensitive data is included in trace payloads by default.Inspect one trace payload for input and output text.RunConfig(trace_include_sensitive_data=False), or set the env var process-wide.
Tool calls never fireThe model does not support function calling, or the endpoint does not pass the tool schema through.Inspect the returned message for a tool_call attribute instead of reading the text.Confirm tool-call support against the endpoint before building on it.
Input tokens grow every stepAgent runs resend the full history on each step.Track input tokens per run as its own metric.Trim history between steps or cap the number of steps.

When this is the wrong move

Cases where the Agents SDK is the wrong tool

SituationWhy it breaksDo this instead
You cannot disable tracing for compliance reasons and have no platform keyTracing is on by default and exports to the OpenAI platform; a missing key means a failure log per run rather than silence.Decide the tracing posture before adopting the SDK, not after.
Your endpoint serves only chat completionsThe default transport is the Responses API; forcing it means fighting a default on every agent.Use a framework that defaults to chat completions, or set the API globally at startup.
You only need one model callThe SDK's value is orchestration — handoffs, guardrails, sessions. A single call carries that overhead for nothing.Use the OpenAI SDK directly.

Where the money actually goes, ranked

Cost advice tends to be a list of tricks with no ordering. These five are ordered by how much they typically move a real bill, and each one names the number that tells you whether it worked.

LeverWhy it worksWhat to measure to know it worked
Switch to a cheaper modelUnit price drops across every token you send.Output quality on your actual task — not on a benchmark.
Cut context per requestInput is re-sent every turn on agent workloads; trimming it compounds.Tokens per task, before and after.
Cache the stable prefixA cached prefix is billed below the uncached rate where the provider supports it.Cache-hit rate in the usage response.
Move non-urgent work to batchBatch tiers trade latency for a lower rate; the work is identical.Nothing changes in quality — only in when it finishes.
Cap output lengthOnly helps when output genuinely dominates, which it does on generation tasks.Whether output is actually the larger half of your bill.

Concretely: one day at 1200k input / 200k output costs $0.240 on gpt-5.6-luna and $0.396 on deepseek-v4-flash — a difference of $0.156 per day, or $3.12 over a 20-day month.

The ordering is not universal — it depends on your token shape. On agent loops, context size outranks model choice; on bulk generation, model choice wins because there is no context to trim. Work out which half of the sum your workload sits on before spending effort on the wrong lever. Our output vs input pricing breakdown covers how to tell.

FAQ

Why does every run log a tracing error after I changed the endpoint?

Because tracing is enabled by default and exports to the OpenAI platform, and it is configured separately from the model client. Pointing the model at a gateway does not move the trace exporter. With no platform key in the environment, each run emits a tracing failure — harmless to your answers, but noisy in your logs and easy to mistake for a routing problem. Either call set_tracing_disabled(True), or keep tracing and give it its own key with set_tracing_export_api_key().

Should I reuse the gateway key for tracing?

No, and the SDK gives you a specific switch for this. Pass use_for_tracing=False when you call set_default_openai_client(), then configure tracing separately. Otherwise the SDK may send your gateway credential to the OpenAI platform's tracing endpoint — a credential-leak shape that has nothing to do with cost but is worth two lines to avoid.

Why do I get 404 on a base URL that works everywhere else?

The SDK defaults to the Responses API, so a plain agent targets POST /v1/responses. An endpoint that serves only /v1/chat/completions returns 404, which reads exactly like a wrong base URL and is not one. Make the transport explicit with OpenAIChatCompletionsModel, or call set_default_openai_api("chat_completions") when the address and key come from environment variables.

How do I keep customer data out of traces?

Two mechanisms, both official. Per run: RunConfig(trace_include_sensitive_data=False). Process-wide: set OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA=0 before the app starts. There are equivalent flags for logs — OPENAI_AGENTS_DONT_LOG_MODEL_DATA and OPENAI_AGENTS_DONT_LOG_TOOL_DATA — which also control whether failures retain payload-bearing diagnostic detail.

Where does the base URL actually go?

On the AsyncOpenAI client, not on the Agent. The agent takes a model object, and the model takes the client. Three scopes are available: set_default_openai_client() globally, a ModelProvider passed through RunConfig per run, or Agent.model per agent — which is what lets different agents in one application use different endpoints.

Get API access