LangChain with a custom endpoint

LangChain wraps the OpenAI client rather than replacing it, so pointing it at a different endpoint is one kwarg. The part that costs money is subtler: changing the base URL also changes whether LangChain asks for token usage on streamed responses, and that decision is what your per-request cost numbers rest on.

In short: LangChain disables the default for stream_usage when a non-OpenAI base URL is set, so usage_metadata can start returning None after an endpoint swap — set stream_usage=True explicitly before trusting any per-request cost number.
You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment
Confirm the key works first. One command, no SDK, costs nothing:
curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.

What you need before you start

Python — the minimal change

Everything else in your chain survives untouched. Prompts, output parsers, retrievers and agent loops sit above the model object, so replacing the model is the whole change.

import os
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="gpt-5.6-luna",                 # exact ID from /v1/models
    api_key=os.environ["GATEWAY_API_KEY"],
    base_url="https://aicomp.ai/v1",
    temperature=0.2,
    max_retries=2,
)

msg = llm.invoke("In one sentence: what is a rate limit?")
print(msg.content)

Two arguments carry the routing. base_url decides where the request goes and must include the /v1 segment, because the client appends the resource path itself. api_key supplies the credential — and it is worth passing explicitly rather than relying on OPENAI_API_KEY, because a stale variable left in a long-lived shell is the most common cause of a 401 that looks like a broken base URL.

TypeScript — a different shape

The JS client does not mirror the Python one. The base URL belongs inside a configuration object, and putting it at the top level of the constructor does nothing.

import { ChatOpenAI } from "@langchain/openai";

// baseURL belongs inside `configuration`, not at the top level.
const llm = new ChatOpenAI({
  model: "gpt-5.6-luna",
  apiKey: process.env.GATEWAY_API_KEY,
  maxRetries: 2,
  configuration: {
    baseURL: "https://aicomp.ai/v1",
  },
});

const msg = await llm.invoke("In one sentence: what is a rate limit?");
console.log(msg.content);

The one behaviour worth testing before you trust your cost numbers

LangChain inspects whether a non-OpenAI base URL is configured when it decides whether to default-enable stream_usage. When one is set, the default is switched off — the stated reason being that many non-OpenAI endpoints do not support streaming token usage. The consequence is specific and easy to miss: a chain that used to report usage_metadata can start reporting None, with no exception, no warning, and a response that looks entirely correct.

That is not a LangChain bug. It is a reasonable default applied to an endpoint whose capabilities it cannot know. What makes it expensive is that per-request cost attribution usually reads exactly that field, so the dashboards quietly go empty while the bill keeps accruing.

import os
from langchain_openai import ChatOpenAI

# The point of this snippet: find out whether the endpoint actually reports
# token usage back. If usage_metadata is None, your cost dashboards are
# blind from here on -- not wrong, just empty.
llm = ChatOpenAI(
    model="gpt-5.6-luna",
    api_key=os.environ["GATEWAY_API_KEY"],
    base_url="https://aicomp.ai/v1",
    stream_usage=True,          # ask for usage on streamed responses
)

msg = llm.invoke("Say: ok")
print("usage_metadata:", getattr(msg, "usage_metadata", None))
print("response_metadata keys:", sorted(getattr(msg, "response_metadata", {}).keys()))

# Non-streaming call: the usage block here is what a cost callback would read.
raw = llm.invoke("Say: ok", config={"callbacks": []})
print("input_tokens :", (raw.usage_metadata or {}).get("input_tokens"))
print("output_tokens:", (raw.usage_metadata or {}).get("output_tokens"))

If the probe prints None, you have three options in increasing order of effort: ask the endpoint whether it can return a usage block on streamed responses and enable it if so; fall back to non-streaming calls for the subset of traffic you need metered; or meter tokens at the call site with your own counter and use the gateway's usage log only for reconciliation. The third is the one that survives an endpoint change.

What it costs

Models reachable from LangChain — gateway rate vs official list
ModelGateway rate
in / out per 1M tokens
Official list
in / out per 1M tokens
Diff
gpt-5.6-luna$0.1 / $0.6$0.2 / $1.250%
claude-sonnet-5$1 / $5$2 / $1050%
deepseek-v4-flash$0.22 / $0.66— / ——

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.

Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget. Across the 216 models in our cleaned catalogue the median output-to-input ratio is 4.0×, which is why the model column matters more than any code change in this page.

What actually goes over the wire

Every OpenAI-compatible call is an HTTP POST to a base URL plus a resource path, with two things that decide everything else: the path, and the bearer token.

POST /v1/chat/completions HTTP/1.1
Host: aicomp.ai
Authorization: Bearer sk-...
Content-Type: application/json

{"model": "gpt-5.6-luna",
 "messages": [{"role": "user", "content": "..."}],
 "stream": true}

Read that literally, because three separate failure modes hide in it. The Host comes from your base URL — if the request arrives somewhere unexpected, this is the line that was wrong. The Authorization header is the only thing identifying you, so a shared key means indistinguishable spend. And the model value inside the body is not validated against anything client-side: send an ID the server does not serve and you get a server-side error, not a client-side warning.

Nothing else in the request is provider-specific. The message array, the tool schemas, the stream flag and the response shape are the wire format, and they are what makes one base URL substitutable for another. This is also the reason the change is reversible: you are not recompiling anything, you are editing one string.

LangChain failures and what each one is actually telling you

SymptomWhat is actually happeningHow to confirmFix
usage_metadata is None after the switchstream_usage is defaulted off when a non-OpenAI base URL is set, so no usage block is requested or returned.Print usage_metadata on one response. None means this is the cause.Pass stream_usage=True explicitly, or meter tokens yourself at the call site.
Requests still go to api.openai.comThe kwarg was not applied — a mixed openai_api_base / base_url call, or a client constructed before the setting existed.Log the resolved base URL from the client object, not from your config file.Pass base_url as a kwarg and print client.base_url once at startup.
401 with a key you just createdA stale OPENAI_API_KEY in the shell or in a .env that loads after your assignment.Print the first six characters of the key the client actually holds.Pass api_key explicitly as a kwarg rather than depending on the environment.
model_not_found for a model the endpoint servesAliases are not resolved by every gateway; concrete dated IDs usually are.Compare the string against the endpoint's /v1/models output.Copy the exact ID instead of reconstructing it from memory.
Reasoning traces vanish on reasoning modelsChatOpenAI targets the official specification and drops non-standard fields such as reasoning_content.Inspect the raw response_metadata for a reasoning key.Use the provider's own LangChain integration package if one exists.
TS app ignores the base URL entirelybaseURL at the top level of the constructor is not read; it belongs under configuration.Send one request and check which host the error or usage log names.Move baseURL into the configuration object.

When this is the wrong move

Cases where plain ChatOpenAI is the wrong abstraction

SituationWhy it breaksDo this instead
You need a provider that extends the response formatChatOpenAI preserves the official OpenAI schema only. Reasoning fields, provider-specific metadata and custom finish reasons are dropped.Use the provider's dedicated LangChain package where one exists.
Your application depends on streaming token countsThe default is off for non-OpenAI endpoints, so anything built on usage_metadata needs an explicit decision rather than a default.Set stream_usage=True and verify it returns data before rolling out.
You are mixing two routes in one processEnvironment variables are read once at init; a client built before the change keeps the old host for its whole lifetime.Construct the client after configuration, and pass routing as kwargs.

Rolling it out without finding out the hard way

The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:

  1. Prove it on one non-critical chain that already has a cost callback. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
  2. Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
  3. Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
  4. Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.

Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.

FAQ

Why does my token usage disappear after I change the base URL?

This is the one LangChain behaviour worth knowing before you wire up cost tracking. LangChain decides whether to default-enable stream_usage by inspecting whether a non-OpenAI base URL is in play — when one is set, the default is switched off, on the assumption that many non-OpenAI endpoints do not support streaming token usage. The practical result is that a chain which reported usage_metadata against OpenAI can start reporting None against a gateway, with no error and no warning. Set stream_usage=True explicitly and read one response back before you trust any per-request cost number. The section above shows the four-line probe.

Should I use base_url or openai_api_base?

base_url is the current name; openai_api_base is the older alias and still resolves in many versions. Pick one and stay with it. Passing both is the failure people actually hit — two kwarg names for one setting means whichever the constructor reads last wins, and the result looks like a base URL that silently reverts to the default.

What is the precedence between the kwarg and the environment variables?

Three levels, first match wins: an explicit base_url kwarg, then OPENAI_API_BASE, which LangChain reads at init, then OPENAI_BASE_URL, which the underlying OpenAI SDK reads. That third one matters more than it looks, because it is the variable LangChain also inspects for the stream_usage decision described above — so setting the environment variable rather than the kwarg can change behaviour you did not intend to change.

Why do I lose the reasoning content from reasoning models?

LangChain's ChatOpenAI targets the official OpenAI specification and does not extract or preserve non-standard response fields such as reasoning_content. If you are using a provider that extends the Chat Completions format, the official guidance is to use that provider's own integration package. The failure mode here is quiet: the answer arrives, the reasoning trace does not, and nothing complains.

Where does baseURL go in the TypeScript version?

Inside the configuration object, not at the top level of the constructor. new ChatOpenAI({ configuration: { baseURL: ... } }). Setting a generic environment variable instead can leave the application pointed at a different default endpoint than you think, because the JS client does not consult the same variables the Python one does.

Get API access