DSPy with an OpenAI-compatible endpoint

DSPy is a framework for optimising prompts, and that is exactly why its bills surprise people. The requests that dominate are not the ones your program makes at runtime — they are the ones compile() makes while searching for a better prompt.

In short: In DSPy the dominant cost is compile(), not inference: an optimiser calls the LM several times per trainset row, so a 200-row MIPROv2 run is roughly a thousand billed calls before the program serves a single request.
You need a key before the code below runs. Create an account, generate a key, and copy the base URL (https://aicomp.ai/v1). Create one free →
Check current rates → Free to sign up · $1 minimum top-up · No prepayment
Confirm the key works first. One command, no SDK, costs nothing:
curl https://aicomp.ai/v1/models \
  -H "Authorization: Bearer sk-your-gateway-key"

A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.

Step 1 — configure the LM

import dspy

lm = dspy.LM(
    "openai/gpt-5.6-luna",
    api_base="https://aicomp.ai/v1",
    api_key="sk-your-gateway-key",
)
dspy.configure(lm=lm)

The openai/ prefix is what selects the OpenAI-compatible route through the provider resolver. A gateway-served model generally needs it.

Step 2 — model_type, when the endpoint needs it

# Endpoints that are not OpenAI's sometimes need the model type spelled out.
lm = dspy.LM(
    "openai/deepseek-v4-flash",
    api_base="https://aicomp.ai/v1",
    api_key="sk-your-gateway-key",
    model_type="chat",
)

Endpoints that are not OpenAI's own sometimes need the model type stated explicitly. The failure it produces tends to read like a model or format problem rather than a configuration one, so it is worth trying early when a call that works with curl fails through the framework.

Step 3 — the part nobody budgets: compile()

An optimiser runs your module across the trainset repeatedly: generating candidates, scoring them against the metric, retrying the ones that fail. Each of those is a request. A single compile() call can be hundreds of them.

import dspy
from dspy.teleprompt import BootstrapFewShot

optimiser = BootstrapFewShot(metric=answer_exact_match)
# 每一次 compile 都会在训练集上反复调用 LM —— 这是账单的主体
compiled = optimiser.compile(student=classifier, trainset=train_set)
One compile() run at 1.5k input / 400 output per call — gateway rates checked 2026-09-26
Optimisergpt-5.6-lunaclaude-sonnet-5deepseek-v4-flash
BootstrapFewShot
50 trainset rows, ~3 calls each — 150 calls
$0.0585$0.525$0.0891
MIPROv2 (light)
50 rows, ~5 calls with candidate search — 250 calls
$0.0975$0.875$0.1485
MIPROv2 (full)
200 rows, ~5 calls — a realistic tuning run — 1000 calls
$0.39$3.5$0.594

The figures are computed from the rate table, at the stated call counts — your own trainset and optimiser settings will differ. The point is the order of magnitude: compiling is not a rounding error next to inference.

Compared with a single inference call at $0.00039, the ratio is the point. If your bill looks inexplicable, this is usually where it went.

Step 4 — count the calls once

Rather than estimating, wrap the LM and count. It is a few lines, and it converts the largest unknown in a DSPy project into a number you can act on.

import dspy

class CountCalls(dspy.Module):
    def __init__(self):
        self.n = 0

# 包装 LM,统计 compile 期间到底打了多少次
class CountingLM(dspy.LM):
    def __init__(self, inner):
        self.inner = inner
        self.calls = 0

    def __call__(self, *a, **kw):
        self.calls += 1
        return self.inner(*a, **kw)

wrapped = CountingLM(lm)
dspy.configure(lm=wrapped)
# ... 跑完 compile 之后
print("calls during compile:", wrapped.calls)

Two levers follow from the count. Optimise against a cheaper model — the optimiser is searching over prompts, and a cheaper model often finds a comparable one — then run the compiled program on the model you intend to ship. And shrink the trainset before shrinking anything else, because cost scales linearly with rows.

DSPy failures, ordered by how much they cost

SymptomWhat is actually happeningHow to confirmFix
Bill far above what inference explainscompile() calls the LM many times per trainset row; that is the dominant cost.Count calls during one compile and multiply by your per-call rate.Optimise on a cheaper model, or reduce the trainset before anything else.
Model not found, or a provider errorThe openai/ prefix is missing, so the resolver looks for a provider it knows.Test the same model string with a raw curl call.Prefix the ID with openai/ and copy it from /v1/models.
Calls fail in ways that look like model errorsThe endpoint needs model_type stated explicitly.Compare a raw call against the framework call with identical credentials.Pass model_type='chat'.
Optimisation quality dropped after switching modelsThe optimiser and the runtime model were changed together, so the cause is ambiguous.Change one at a time and re-measure on your own metric.Optimise cheap, run on the model you ship, and verify with your metric.
Requests never reach the endpointdspy.configure was not called, or was called with a different object than the one you tested.Check the usage log for the configured model ID.Configure once at startup and verify with a single small program.

When this is the wrong move

Where the optimisation loop costs more than it returns

SituationWhy it breaksDo this instead
Your trainset is large and the metric is noisyCost scales with rows, and a noisy metric makes extra candidates worthless.Cut the trainset and tighten the metric before tuning.
You are optimising and serving on the same expensive modelThat doubles the most expensive part of the pipeline for little gain.Optimise on a cheap model; serve on the one you intend to ship.
You have not counted calls per compileWithout it, every cost decision here is a guess.Wrap the LM and count once.

Rolling it out without finding out the hard way

The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:

  1. Prove it on one compile against a 20-row subset. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
  2. Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
  3. Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
  4. Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.

Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.

Where the money actually goes, ranked

Cost advice tends to be a list of tricks with no ordering. These five are ordered by how much they typically move a real bill, and each one names the number that tells you whether it worked.

LeverWhy it worksWhat to measure to know it worked
Switch to a cheaper modelUnit price drops across every token you send.Output quality on your actual task — not on a benchmark.
Cut context per requestInput is re-sent every turn on agent workloads; trimming it compounds.Tokens per task, before and after.
Cache the stable prefixA cached prefix is billed below the uncached rate where the provider supports it.Cache-hit rate in the usage response.
Move non-urgent work to batchBatch tiers trade latency for a lower rate; the work is identical.Nothing changes in quality — only in when it finishes.
Cap output lengthOnly helps when output genuinely dominates, which it does on generation tasks.Whether output is actually the larger half of your bill.

Concretely: one day at 1500k input / 400k output costs $0.390 on gpt-5.6-luna and $0.594 on deepseek-v4-flash — a difference of $0.204 per day, or $4.08 over a 20-day month.

The ordering is not universal — it depends on your token shape. On agent loops, context size outranks model choice; on bulk generation, model choice wins because there is no context to trim. Work out which half of the sum your workload sits on before spending effort on the wrong lever. Our output vs input pricing breakdown covers how to tell.

FAQ

Why is my DSPy bill so much higher than inference would suggest?

Because the optimiser, not the program, is doing most of the calling. compile() runs your module over the trainset repeatedly — generating, scoring, retrying — and each of those is a billed request. A tuning run that feels like one operation can be hundreds of calls. Measure it once by counting calls during compile, and the number usually explains the invoice immediately.

Do I need the openai/ prefix?

Yes for an OpenAI-compatible endpoint. DSPy routes the model string through its provider resolver, and the openai/ prefix is what selects the OpenAI-compatible path. Without it the resolver looks for a provider it knows, and a gateway-served model is generally not one of them.

When is model_type='chat' needed?

With endpoints that do not advertise themselves as OpenAI's own. The symptom is a failure that reads like a model or format problem rather than a configuration one, which is what makes it slow to diagnose. If a call that works in curl fails through DSPy with the same key and base URL, this is worth trying early.

How do I find out how many calls compile() made?

Wrap the LM object and increment a counter on each call, then configure DSPy with the wrapper and read the counter after compiling. It is a few lines and it replaces a guess with a number. Do it once per optimiser setting — the counts differ substantially between them.

Is it cheaper to optimise against a cheap model?

Often yes, and it is the most effective single lever on this framework. The optimiser is searching over prompts, not evaluating final quality, so a cheaper model frequently finds a comparable prompt. Optimise on the cheap model and run the compiled program on the model you intend to ship. Verify on your own metric — do not assume it transfers.

Related

Get API access