LiteLLM retries and timeouts: the budget you did not set
Retries are the only setting on this page that can silently multiply your bill. LiteLLM has two retry knobs with confusingly similar names, a fallback chain that re-runs the retry loop at every hop, and a per-error policy whose default behaviour retries errors that can never succeed. None of it warns you.
https://aicomp.ai/v1).
Create one free →
curl https://aicomp.ai/v1/models \
-H "Authorization: Bearer sk-your-gateway-key"
A JSON list of model IDs means the key is good. Invalid token means it was copied wrong.
Two knobs, and only one of them is yours
LiteLLM's own documentation draws the distinction explicitly: num_retries is LiteLLM's retry loop, while max_retries is the provider SDK's internal retry count. For a request that goes through the router, LiteLLM takes ownership of retries and pins the provider client to max_retries: 0. The stated reason is worth quoting in full, because it is the whole argument of this page — it is what stops a deployment's num_retries: N from being applied twice and turning one request into (1 + N)² upstream calls.
So the rule is simple: set num_retries, and treat max_retries as the router's business. A max_retries you place in the request body or in litellm_params has no effect on a proxy request, which surprises people who set both and then cannot work out why doubling one of them changed nothing.
Where the value can come from, and what wins
num_retries can be set in four places. Ranked highest first: the x-litellm-num-retries request header (proxy only), num_retries in the request body, num_retries in a deployment's litellm_params, and num_retries in litellm_settings as the router-wide default. The useful consequence is that a caller can always override the retry count for one request — including setting it to zero to disable retries entirely — regardless of what the deployment or the global default says.
# litellm_config.yaml
model_list:
- model_name: primary
litellm_params:
model: openai/aicomp.ai-model-id
api_base: https://aicomp.ai/v1
api_key: os.environ/GATEWAY_API_KEY
rpm: 200
- model_name: primary
litellm_params:
model: openai/cheaper-model-id
api_base: https://aicomp.ai/v1
api_key: os.environ/GATEWAY_API_KEY
rpm: 200
router_settings:
num_retries: 2
timeout: 30
# Cap the depth of the fallback chain. Every hop re-runs the retry loop,
# so upstream calls are bounded by (1 + num_retries) * (1 + fallback depth).
fallbacks: [{"primary": ["primary"]}]
retry_policy:
# Errors that will never succeed on a retry. Retrying them is pure spend.
AuthenticationErrorRetries: 0
NotFoundErrorRetries: 0
ContentPolicyViolationErrorRetries: 0
BadRequestErrorRetries: 1
# Errors where a retry is genuinely worth it.
RateLimitErrorRetries: 3
TimeoutErrorRetries: 2
InternalServerErrorRetries: 2
ServiceUnavailableErrorRetries: 2
DefaultRetries: 1
The policy block is the part with real money behind it. Fields resolve from most specific to least specific, and any field you leave unset defers to the next one down. That has a consequence the documentation states plainly: a policy that only sets DefaultRetries retries 404s too.
What a retry setting actually costs
Two numbers matter, and they are not the same number. The first is wall clock — how long a worker is held. LiteLLM backs off with min(1 × 2ⁿ + jitter, 10) seconds, with jitter of up to one second, so a single wait never exceeds about eleven seconds but three of them add up quickly. The second is upstream calls — how many billable requests one user request can produce. Because every fallback hop re-runs the retry loop, the ceiling is (1 + num_retries) × (1 + fallback depth).
| Setting | Backoff added min(1×2ⁿ, 10) + jitter |
Upstream calls with 1 fallback | Calls per 1,000 user requests |
|---|---|---|---|
num_retries=0 | 0–0s | 2 | 6 |
num_retries=1 | 1–2s | 4 | 12 |
num_retries=2 | 3–5s | 6 | 18 |
num_retries=3 | 7–10s | 8 | 24 |
num_retries=4 | 15–19s | 10 | 30 |
num_retries=5 | 25–30s | 12 | 36 |
Backoff column computed from LiteLLM's documented constants (initial delay 1s, max delay 10s, jitter up to 1s). The upstream-call column is the formula above with one fallback hop. Neither number is measured against our own traffic — they are what the configuration implies.
Counting what actually leaves the process
Every number above is what the configuration implies. The only way to know what a live process really does is to count it. A callback logger gives you that in about fifteen lines, and it is worth running once against production-shaped traffic before you trust any retry setting.
import litellm, os, time
# Count what actually leaves the process. Retry settings are a belief until
# you have counted real upstream attempts.
class Counter(litellm.integrations.custom_logger.CustomLogger):
def __init__(self):
self.attempts = {}
async def async_log_success_event(self, kwargs, response_obj, start_time, end_time):
self._tick(kwargs)
async def async_log_failure_event(self, kwargs, response_obj, start_time, end_time):
self._tick(kwargs)
def _tick(self, kwargs):
m = kwargs.get("model", "?")
self.attempts[m] = self.attempts.get(m, 0) + 1
print(f"attempt #{self.attempts[m]} -> {m}")
counter = Counter()
litellm.callbacks = [counter]
t0 = time.time()
try:
r = litellm.completion(
model="primary",
messages=[{"role": "user", "content": "Say: ok"}],
num_retries=2,
timeout=30,
)
print("ok:", r.choices[0].message.content)
except Exception as e:
print("raised:", type(e).__name__, str(e)[:200])
print(f"\nwall clock: {time.time() - t0:.1f}s")
print("upstream attempts by model:", counter.attempts)
Two things to read off the output. The per-model attempt count tells you whether the fallback chain is doing what you think, and the wall-clock number tells you how long a worker actually sat there. If the attempt count is higher than your policy suggests, the usual cause is a retry loop you did not know was running — most often a second one inside the calling application.
# What a retry setting actually costs you, in seconds and in upstream calls.
# Constants from LiteLLM: initial delay 1s, max delay 10s, jitter up to 1s.
# Delay for attempt n = min(1 * 2**n + jitter, 10)
INITIAL, MAX_DELAY, JITTER = 1, 10, 1
def worst_case_seconds(num_retries):
total = 0.0
for n in range(num_retries):
total += min(INITIAL * (2 ** n), MAX_DELAY) + JITTER # jitter at its worst
return total
def best_case_seconds(num_retries):
total = 0.0
for n in range(num_retries):
total += min(INITIAL * (2 ** n), MAX_DELAY) + 0.0 # jitter at zero
return total
def upstream_call_ceiling(num_retries, fallback_depth):
"""Upper bound on billable upstream calls for ONE user request."""
return (1 + num_retries) * (1 + fallback_depth)
for n in range(6):
print(f"num_retries={n}: backoff {best_case_seconds(n):5.1f}-{worst_case_seconds(n):5.1f}s "
f"| ceiling with 1 fallback: {upstream_call_ceiling(n, 1):2d} upstream calls")
Errors that can never succeed on a retry
LiteLLM normalises provider errors into OpenAI-compatible exception types, and the mapping marks several as non-retryable: authentication errors, permission errors, context-window overflow, and content-policy violations. A retry cannot change the outcome for any of them, because the request being resent is identical to the one that was rejected.
The other side of that line is where retries earn their keep — rate limits, timeouts, connection errors, and 5xx responses. Splitting your policy along exactly this boundary is most of the value in writing one. Everything else is tuning.
What it costs when a retry burns tokens
| Model | Gateway rate in / out per 1M tokens | Official list in / out per 1M tokens | Diff |
|---|---|---|---|
| gpt-5.6-luna | $0.1 / $0.6 | $0.2 / $1.2 | 50% |
| claude-sonnet-5 | $1 / $5 | $2 / $10 | 50% |
| deepseek-v4-flash | $0.22 / $0.66 | — / — | — |
Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget.
Rates checked 2026-09-26. Gateway rates move with upstream promotions — verify the current number in your dashboard before committing to a budget. Across the 216 models in our cleaned catalogue the median output-to-input ratio is 4.0×.
Retry failures and what each one is actually telling you
| Symptom | What is actually happening | How to confirm | Fix |
|---|---|---|---|
| The bill grows but request volume is flat | Retries and fallback hops are billed as real requests, and failed ones still consumed input tokens. | Count upstream attempts per user request with a callback logger. | Cap fallback depth, and set non-retryable error types to zero. |
| 404s are being retried | A RetryPolicy that only sets DefaultRetries defers to it for every error, including not-found. | Watch the attempt count on a request you know 404s. | Set NotFoundErrorRetries: 0 explicitly. |
| Setting max_retries changed nothing | The router pins the provider client to max_retries: 0 and owns retries itself. | Compare behaviour with num_retries instead. | Use num_retries for proxy requests; treat max_retries as router-owned. |
| A request holds a worker for minutes | Each retry adds up to about eleven seconds of backoff, and the per-attempt timeout applies to every attempt. | Time a request you know will fail all the way through. | Budget timeout and num_retries together; lower the retry count on expensive models. |
| Retries stop working after moving to the proxy | The header, body, deployment and global levels resolve differently depending on whether you are on the SDK or the proxy path. | Send one request with an explicit num_retries and count attempts. | Set it at the level your deployment actually reads, and verify by counting. |
| Cost spikes during an incident rather than after it | Fallbacks route traffic to whatever is next in the chain; if the primary is rate-limited, the fallback absorbs the same volume. | Plot upstream calls per minute by model during the incident. | Make the fallback cheaper than the primary, and cap the chain depth. |
When this is the wrong move
Cases where more retries make things worse
| Situation | Why it breaks | Do this instead |
|---|---|---|
| You are retrying to buy throughput | Retries consume the same rate-limit window that just refused you, so they convert throttling into a queue that hides the real ceiling. | Reduce concurrency and add it back while watching the error rate. |
| The error is an authentication or permission error | Both are marked non-retryable upstream. Resending the identical request gets the identical rejection, billed twice. | Fix the credential; set AuthenticationErrorRetries to zero. |
| You have a deep fallback chain and a high retry count | The two multiply. The ceiling is (1 + num_retries) × (1 + fallback depth) upstream calls for one user request. | Cap the chain first, then tune retries. |
| You are retrying inside the calling application as well | Two independent retry loops multiply, and neither can see the other's budget. | Pick one layer to own retries — usually the one closest to the provider. |
Rolling it out without finding out the hard way
The failure mode you want to avoid is discovering a problem through a production bill or a customer-visible error. Four steps, in order:
- Prove it on one non-critical route with a callback logger attached. One call, from a script, with an explicit timeout and the model ID echoed back. You are testing reachability, authentication and model availability — three things that can fail independently.
- Measure before you switch. Record tokens per task and cost per task on the current path first. Without that baseline, "it got cheaper" is an impression, not a result.
- Move one workload, not everything. Pick the workload with the most predictable shape — batch jobs over interactive traffic. Leave the interactive path on the old configuration until the batch numbers are in.
- Decide the rollback condition in advance. Write down what makes you revert (error rate above X, cost per task above Y, latency above Z) before you start, so the decision is not made under pressure.
Keep the base URL in configuration, never inline. That single choice is what makes step four take a minute instead of an afternoon.
FAQ
Are <code>num_retries</code> and <code>max_retries</code> the same setting?
No, and conflating them is the most expensive mistake in this area. num_retries is LiteLLM's own retry loop; max_retries is the provider SDK's internal retry count. LiteLLM's own documentation states that for a call going through the router it pins the provider client to max_retries: 0, and that this is deliberate — it is what stops a num_retries: N deployment value from being applied twice and turning one request into (1 + N)² upstream calls. The practical rule: control retries with num_retries, and treat max_retries as a setting the router owns.
Why is my 404 being retried?
Because RetryPolicy resolves from the most specific field to the least specific, and any field you leave unset defers to the next one. A policy that only sets DefaultRetries therefore retries 404s as well — a deleted response ID or an unknown deployment name gets the full retry budget. LiteLLM's documentation calls this out directly and gives the fix: set NotFoundErrorRetries: 0 explicitly. The same logic applies to authentication errors, which is why AuthenticationErrorRetries: 0 belongs in every policy.
How many upstream calls can one request actually trigger?
The upper bound is (1 + num_retries) × (1 + fallback depth), because each fallback hop runs the retry loop again from the start. With num_retries: 2 and one fallback, a single user request can bill six upstream calls — and every one of them is a real request that consumed input tokens, including the ones that failed. The table above is generated from that formula rather than written by hand.
Which errors should never be retried?
LiteLLM's exception mapping marks several as non-retryable: authentication and permission errors, context-window overflow, and content-policy violations. A retry cannot fix any of them — the request is the same request, so the answer is the same answer. What a retry can fix is the transient class: rate limits, timeouts, connection errors, and 5xx responses. Splitting your policy along that line is most of the value in the exercise.
How long can a retried request block a worker?
LiteLLM backs off with min(1 × 2^n + jitter, 10) seconds, where n is the attempt number, and the jitter is up to one second. That caps a single wait at 10–11 seconds, so num_retries: 3 adds roughly 7–10 seconds of waiting on top of however long each attempt takes. Against a 30-second per-attempt timeout, a fully exhausted retry chain can hold a worker for well over a minute — which is why the timeout and the retry count have to be budgeted together, not separately.
Does a retry cost money if it fails?
It can. A request that reaches the model and then times out has already been billed for its input tokens. This is the failure mode where costs grow fastest, because nothing in the application signals that money was spent — the caller only sees an exception. Counting attempts per model, as in the probe above, is the only way to see it.