Troubleshooting: LLM timeouts and upstream failures
An AI agent run failed, or a conversational-IVR menu stopped understanding callers. Four error codes cover the path between your agent and the language model:LLM_TIMEOUT(504) — the language-model round-trip exceeded its deadline before a reply arrived.LLM_UPSTREAM_ERROR(502) — the language-model provider returned an error, or the transport to it broke.IVR_CLASSIFY_LLM_ERROR(503) — the intent classifier that routes a conversational IVR menu call failed on its model call.IVR_SLOT_EXTRACT_LLM_ERROR(503) — the slot-filler that pulls entities (dates, account numbers, names) out of a caller’s utterance failed on its model call.
AGENT_ERROR, LLM_PROVIDER_ERROR,
INVALID_MODEL, and friends), see
Agent runtime errors. This page is the
deeper runbook for the four codes above.
How LLM failures surface on an agent run
A run that dies on a language-model call closes withstatus: "failed"
and an error envelope carrying the code, a message, and a request_id
under meta:
request_id is the one handle that survives the failure — every
language-model call logs it server-side. Quote it when you escalate,
and use it to correlate a failed run with a specific provider blip in
your own logs.
In the dashboard the same failure shows up as an agent run failed
row on the run list, with the code visible on the run detail panel. For
the full set of places a run-level failure appears, see
the Agent run lifecycle.
Retry policy per code
Pick the retry strategy by code — getting this wrong is what turns a provider blip into an incident.
Two rules that cut across all four:
- Back off exponentially and cap attempts. A burst of identical retries against an overloaded provider makes the outage worse. Retry at most two or three times per turn; beyond that, fail the run cleanly and alert.
- Use the
request_id, not the message text. The message is a fixed, customer-safe string — it never echoes the raw provider error (by design; the provider detail is logged server-side). Escalation without therequest_idforces whoever investigates to dig for it. Include it every time.
The IVR NLU path, mapped
Conversational IVR replaces a DTMF menu (“press 1 for balance”) with two model calls per turn:- Classify — decides which intent (menu destination) the caller’s
utterance maps to. Failure →
IVR_CLASSIFY_LLM_ERROR. - Extract slots — pulls entities out of the utterance (the account
number the caller spoke, the date they want). Failure →
IVR_SLOT_EXTRACT_LLM_ERROR.
- Re-ask the caller. In most flows the cheapest recovery is to re-prompt for the same utterance (“Sorry — could you say that again?”). For slot extraction, re-ask for that specific field.
- Fall back to DTMF. If your flow has a DTMF mirror for the same menu (a “press 1” fallback), route to it after the first NLU 503. This keeps callers moving during a provider outage.
- Fail the flow to a queue. For flows where a wrong turn is worse than a wait (payments, account changes), drop the call into a queue with a human rather than guessing.
request_id.
What NOT to do
- Do not blind-replay the same envelope on
LLM_UPSTREAM_ERROR. A 502 means the provider already rejected the request — replaying the identical payload with no backoff hammers an already-failing upstream and burns your cost ceiling. Retry once or twice with backoff; then stop and escalate. - Do not re-run the whole turn without the
request_id. Escalating “it failed, please investigate” with norequest_idmakes the failure un-actionable — quotes from a failed run without its id cannot be traced to a single provider event. Capture the id from themetablock before you do anything else. - Do not treat
LLM_TIMEOUTas deterministic. A timeout is usually a slow or stalled call, not a permanent rejection — a retry with backoff is expected to succeed. Reserve escalation for timeouts that survive two or three backoff attempts. - Do not parse the
messagetext programmatically. The message is a customer-safe string with no provider detail — branch onerror.code, not the message.
When to escalate
Escalate with therequest_id (and the flow name for IVR) when any of
these holds:
- The same
LLM_UPSTREAM_ERRORrepeats across more than three backoff attempts on a single run. LLM_TIMEOUTsurvives two or three retries on prompts that used to succeed.IVR_CLASSIFY_LLM_ERRORorIVR_SLOT_EXTRACT_LLM_ERRORhits more than a handful of calls on the same flow in a short window — treat it as a provider outage and flip the flow to DTMF first.- The status page shows a provider
incident — quote the
request_idanyway; it still pinpoints the affected calls.