Streaming through a custom endpoint: the usage chunk you are not getting
Turn on stream: true and the token counts you were relying on quietly disappear. Every chunk reports usage: null except one, that one has an empty choices array, and if the stream breaks you never receive it at all. None of this raises an error, which is what makes it expensive.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
What streaming does to the usage field
OpenAI's API reference is explicit about the shape. With stream_options set to include usage, an additional chunk is streamed before the data: [DONE] message. On that chunk, usage carries the statistics for the whole request and choices is always an empty array. On every other chunk, usage is present but null.
Two consequences follow directly, and both are the kind that show up as a bug report rather than as an exception:
- Without the flag there is no count at all. Not a partial count — no
usageobject to read. A per-request cost line that readsusage.prompt_tokensrecords nothing. - With the flag, the useful chunk breaks naive code. It is the only chunk with an empty
choiceslist, so the line that readschoices[0].delta.contentraises on exactly the chunk carrying the numbers.
from openai import OpenAI
client = OpenAI(base_url="https://aicomp.ai/v1", api_key=KEY)
stream = client.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "Summarise this file in three lines."}],
stream=True,
stream_options={"include_usage": True}, # <- produces the final usage chunk
)
text, usage = [], None
for chunk in stream:
if not chunk.choices:
# This is the usage chunk: choices is an empty list, usage is populated.
usage = chunk.usage
continue
delta = chunk.choices[0].delta
if delta.content:
text.append(delta.content)
print("".join(text))
print("usage:", usage)
# -> None means the endpoint ignored stream_options. No error is raised.
# Your token counter is now zero for every request it sends.
That guard — if not chunk.choices — is the whole fix for the second problem. It is also the only place you can distinguish "the endpoint returned no usage" from "the request produced no tokens".
The flag is a request, not a guarantee
This is the part that costs money. stream_options is a field you send; whether the endpoint implements it is up to the endpoint. One that does not will accept the request, stream perfectly good text, and never emit the usage chunk. No error surfaces. Your counter reads null and your dashboard reads zero.
So the check is not "did I set the flag" but "did a usage object come back". Run the snippet above once per endpoint and per model — behaviour can differ between models behind the same base URL — and record the answer. If it is null, you have two honest options: count tokens client-side, or stop streaming for the requests you need to attribute.
A missing chunk is unknown, not zero
OpenAI's documentation carries a note that deserves more attention than it gets: if the stream is interrupted, you may not receive the final usage chunk, and not receiving it does not mean the request consumed no tokens. A client disconnect, a timeout, a dropped connection mid-generation — all of them produce a request that generated output and a tracker that recorded none.
The correct representation in your data model is therefore not 0 but unknown, and unknowns should be reconciled against the provider's own usage view rather than left as zeros in a sum.
What the gap is worth
Below: one day of 2,000 streamed requests at 8k input and 700 output each — a modest assistant-shaped workload, not a heavy one. The first column is what the work actually costs; the last is what a tracker that never receives a usage chunk records.
| Model | Actual cost / day | Actual cost / 20 days | Recorded if usage never arrives |
|---|---|---|---|
gpt-5.6-luna | $2.44 | $48.8 | $0 |
gemini-3.7-flash | $8.625 | $172.5 | $0 |
deepseek-v4-flash | $4.444 | $88.88 | $0 |
claude-sonnet-5 | $23 | $460 | $0 |
Rates checked 2026-09-26; gateway rates move with upstream promotions. The workload shape is illustrative — substitute your own request count and token shape.
# How much spend a missing usage chunk actually hides.
# Shape: 2,000 streamed requests a day, 8k input / 700 output each.
REQUESTS, TIN, TOUT = 2000, 8_000, 700
# $ per 1M tokens -- replace with the rates for the model you route to.
IN_RATE, OUT_RATE = 1.0, 5.0
def day_cost(recorded_fraction):
"""recorded_fraction = share of requests whose usage you actually captured."""
return (REQUESTS * TIN * recorded_fraction) / 1e6 * IN_RATE \
+ (REQUESTS * TOUT * recorded_fraction) / 1e6 * OUT_RATE
for label, f in [("non-streamed (usage always present)", 1.0),
("streamed, stream_options ignored", 0.0),
("streamed, 4% of streams interrupted", 0.96)]:
print(f"{label:40} ${day_cost(f):7.2f}/day ${day_cost(f) * 20:8.2f}/20d")
Two streaming failures that are not about usage
Worth separating, because both get reported as "the endpoint is broken":
- Implemented
/chat/completionsbut notstream: true. A plaincurlreturns a clean response and the client hangs, because clients default to streaming. Always test with streaming on before concluding anything. - Buffered responses. An intermediary that buffers the stream produces no tokens until the end, which reads as a hang and trips client-side timeouts that a non-streamed call would pass. Our LangChain page covers the related case where a non-OpenAI base URL turns streaming usage off inside the framework.
Related guides
- One endpoint for every tool — the hub for tool-specific setup
- OpenAI SDK with a custom endpoint — the non-streamed baseline
- Monitoring API spend — reconciling what you record against what you are billed
- Output vs input pricing — which half of the sum your workload lives on
- LangChain on a custom endpoint — stream_usage being disabled silently
- Endpoint error codes — the failures that return a 200
Streaming-accounting failures and what each one is actually telling you
| Symptom | What is actually happening | How to confirm | Fix |
|---|---|---|---|
| Cost per request is 0.00 for everything | No usage chunk is arriving, so the counter has nothing to read. | Print the usage object for one streamed request. | Set stream_options with include_usage; if still null, count tokens client-side. |
| IndexError at the end of an otherwise working stream | The usage chunk has an empty choices list. | Log the chunk count and the last chunk's choices field. | Guard on empty choices and read usage from that chunk. |
| The client hangs but curl works | curl defaults to non-streaming; clients default to stream: true, and the endpoint may not implement it. | Send the same request with stream: true in curl. | Either use a streaming-capable endpoint or configure the client not to stream. |
| Streamed totals are lower than the provider's usage view | Interrupted streams never deliver the final usage chunk, and those tokens were still generated. | Compare your recorded total against the provider's for the same window. | Record unknown rather than zero; reconcile monthly. |
| The flag works on one model and not another | stream_options is implemented per route or per model on some gateways. | Test the flag once per model, not once per endpoint. | Track capability per model, and fall back to non-streamed where it is absent. |
| Token counts drift from what the dashboard shows | Client-side estimates differ from server-side tokenisation, especially with tool schemas and non-Latin text. | Compare one known request's server usage against your estimate. | Use the server's usage object as truth, never an estimate. |
When this is the wrong move
Four habits that make a streaming integration quietly expensive
| Situation | Why it breaks | Do this instead |
|---|---|---|
| You are about to estimate tokens from character counts | Client-side estimates diverge from server tokenisation, and the server's number is the one you are billed on. | Read the usage object; estimate only for pre-flight guards. |
| You are about to record an interrupted stream as zero | Tokens generated before the interruption were generated and billed. | Record unknown, then reconcile against the provider's usage view. |
| You are about to assume the flag worked because no error appeared | Ignoring an unimplemented request field is silent by design. | Print the usage object once per model and keep the answer. |
| You are streaming everything because it feels faster | Streaming improves time-to-first-token, not throughput, and it is what makes usage accounting fragile. | Stream for interactive output; use non-streamed calls where you need the counts. |
Rolling it out without finding out the hard way
The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:
- Prove it on one streamed request per model, with the usage object printed back. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
- Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
- Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
- Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.
Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.
FAQ
Why is <code>usage</code> null on every chunk?
Because that is the documented behaviour without the flag. OpenAI's API reference states that when you set stream_options with include_usage, one additional chunk is streamed before the data: [DONE] message; the usage field on that chunk holds the statistics for the entire request, the choices field on it is always an empty array, and every other chunk carries usage with a null value. So nulls are not a bug in your endpoint — they are the default.
Why does my accumulation code crash on the last chunk?
Because the chunk that carries the usage has choices: []. Any line that reads chunk.choices[0] unconditionally raises an IndexError on precisely the one chunk you need. Guard on if not chunk.choices: and read usage there, then continue — the text has already been delivered by the earlier chunks.
I set the flag and usage is still null. What now?
stream_options is a field in your request, not a capability the endpoint is obliged to implement. An endpoint that does not support it ignores it silently: no error, no warning, and usage stays null. Confirm it once with a single request that prints the object, and if it comes back null, either count tokens client-side or stop streaming for the requests whose cost you need to attribute.
If the stream is interrupted, did that request cost anything?
Almost certainly yes, and this is the sentence worth taping to a monitor: OpenAI's documentation notes that if the stream is interrupted you may not receive the final usage chunk, and that not receiving it does not mean no tokens were consumed. Tokens generated before the interruption were generated. Treat an interrupted stream as an unknown quantity, not a free one.
Does streaming cost less or more than a normal request?
The same — the tokens are identical and so is the rate. What changes is your ability to measure them, and your exposure to partial work: a stream that dies at 90% still produced 90% of the output tokens. Streaming also changes failure shape rather than cost — a client that hangs on an endpoint that does not emit streaming frames looks identical to one that is merely slow.
How much spend can a missing count actually hide?
All of it, in the worst case. The table above prices one day of 2,000 streamed requests at 8k input and 700 output: a tracker that records zero for every request understates the month by the full amount in the first column. The figure is not a rounding error, and nothing in the application signals that it is happening.