> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Troubleshooting: LLM timeouts and upstream failures on agent runs and IVR NLU

> Read the four LLM-call failure codes — LLM_TIMEOUT, LLM_UPSTREAM_ERROR, IVR_CLASSIFY_LLM_ERROR, IVR_SLOT_EXTRACT_LLM_ERROR — map each to its cause, pick the right retry policy, and know when to stop retrying and escalate with the request id.

# Troubleshooting: LLM timeouts and upstream failures

An AI agent run failed, or a conversational-IVR menu stopped understanding
callers. Four error codes cover the path between your agent and the
language model:

* **`LLM_TIMEOUT`** (504) — the language-model round-trip exceeded its
  deadline before a reply arrived.
* **`LLM_UPSTREAM_ERROR`** (502) — the language-model provider returned an
  error, or the transport to it broke.
* **`IVR_CLASSIFY_LLM_ERROR`** (503) — the intent classifier that routes a
  conversational IVR menu call failed on its model call.
* **`IVR_SLOT_EXTRACT_LLM_ERROR`** (503) — the slot-filler that pulls
  entities (dates, account numbers, names) out of a caller's utterance
  failed on its model call.

For the general agent error surface (`AGENT_ERROR`, `LLM_PROVIDER_ERROR`,
`INVALID_MODEL`, and friends), see
[Agent runtime errors](/troubleshooting/agent-errors). This page is the
deeper runbook for the four codes above.

## How LLM failures surface on an agent run

A run that dies on a language-model call closes with `status: "failed"`
and an `error` envelope carrying the code, a message, and a `request_id`
under `meta`:

```json theme={null}
{
  "status": "failed",
  "error": {
    "code": "LLM_UPSTREAM_ERROR",
    "message": "The language-model provider returned an error."
  },
  "meta": {
    "request_id": "req_01J0A1BC3D"
  }
}
```

The `request_id` is the one handle that survives the failure — every
language-model call logs it server-side. Quote it when you escalate,
and use it to correlate a failed run with a specific provider blip in
your own logs.

In the dashboard the same failure shows up as an **agent run failed**
row on the run list, with the code visible on the run detail panel. For
the full set of places a run-level failure appears, see
[the Agent run lifecycle](/concepts/agent-run-lifecycle).

## Retry policy per code

Pick the retry strategy by code — getting this wrong is what turns a
provider blip into an incident.

| Code | HTTP class | What happened | What to do |
| - | - | - | - |
| `LLM_TIMEOUT` | 504 | The model call hit its abort deadline — usually a long prompt, a burst of provider slowness, or a stalled stream. | **Safe to retry.** Re-issue the same request with exponential backoff (for example 1s, 2s, 4s). If it keeps timing out, shorten the prompt or trim the context the run loads. |
| `LLM_UPSTREAM_ERROR` | 502 | The provider returned an error envelope (overload, internal fault, transport failure) rather than a timeout. | **Short-circuit, then escalate.** A retry is fine — once, with backoff — but if it repeats, do not keep replaying the same envelope. Check the [status page](https://status.orbit.devotel.io), then escalate with the `request_id`. |
| `IVR_CLASSIFY_LLM_ERROR` | 503 | The conversational-IVR intent classifier's model call failed; the call fell back to the menu's no-match path. | **Retry the call, not the envelope.** The caller can usually just re-speak the utterance. If every caller's menu fails classifying, treat it as a provider outage and route the IVR to a fallback DTMF menu or a queue. |
| `IVR_SLOT_EXTRACT_LLM_ERROR` | 503 | The slot-filler that extracts entities (an account number, a date, a name) from the caller's utterance failed. | **Same shape as the classifier** — re-ask the caller for the slot, or fall back to a DTMF/keypad prompt for that field. |

Two rules that cut across all four:

1. **Back off exponentially and cap attempts.** A burst of identical
   retries against an overloaded provider makes the outage worse.
   Retry at most two or three times per turn; beyond that, fail the
   run cleanly and alert.
2. **Use the `request_id`, not the message text.** The message is a
   fixed, customer-safe string — it never echoes the raw provider
   error (by design; the provider detail is logged server-side).
   Escalation without the `request_id` forces whoever investigates to
   dig for it. Include it every time.

## The IVR NLU path, mapped

Conversational IVR replaces a DTMF menu ("press 1 for balance") with
two model calls per turn:

1. **Classify** — decides which intent (menu destination) the caller's
   utterance maps to. Failure → `IVR_CLASSIFY_LLM_ERROR`.
2. **Extract slots** — pulls entities out of the utterance (the account
   number the caller spoke, the date they want). Failure →
   `IVR_SLOT_EXTRACT_LLM_ERROR`.

A 503 from either stage means the model leg of that menu turn failed —
not the telephony leg. The call is still up; the NLU layer is what
broke. Your recovery choices are per-flow:

* **Re-ask the caller.** In most flows the cheapest recovery is to
  re-prompt for the same utterance ("Sorry — could you say that
  again?"). For slot extraction, re-ask for that specific field.
* **Fall back to DTMF.** If your flow has a DTMF mirror for the same
  menu (a "press 1" fallback), route to it after the first NLU 503.
  This keeps callers moving during a provider outage.
* **Fail the flow to a queue.** For flows where a wrong turn is worse
  than a wait (payments, account changes), drop the call into a queue
  with a human rather than guessing.

Which of these is right is a per-flow decision you own — the platform
surfaces the 503 and you decide the recovery path. If every caller on
the same flow is failing, treat it as an outage, flip the flow to its
DTMF mirror, and escalate with the `request_id`.

## What NOT to do

* **Do not blind-replay the same envelope on `LLM_UPSTREAM_ERROR`.**
  A 502 means the provider already rejected the request — replaying the
  identical payload with no backoff hammers an already-failing upstream
  and burns your cost ceiling. Retry once or twice with backoff; then
  stop and escalate.
* **Do not re-run the whole turn without the `request_id`.** Escalating
  "it failed, please investigate" with no `request_id` makes the failure
  un-actionable — quotes from a failed run without its id cannot be
  traced to a single provider event. Capture the id from the `meta`
  block before you do anything else.
* **Do not treat `LLM_TIMEOUT` as deterministic.** A timeout is usually
  a slow or stalled call, not a permanent rejection — a retry with
  backoff is expected to succeed. Reserve escalation for timeouts that
  survive two or three backoff attempts.
* **Do not parse the `message` text programmatically.** The message is a
  customer-safe string with no provider detail — branch on `error.code`,
  not the message.

## When to escalate

Escalate with the `request_id` (and the flow name for IVR) when any of
these holds:

* The same `LLM_UPSTREAM_ERROR` repeats across more than three backoff
  attempts on a single run.
* `LLM_TIMEOUT` survives two or three retries on prompts that used to
  succeed.
* `IVR_CLASSIFY_LLM_ERROR` or `IVR_SLOT_EXTRACT_LLM_ERROR` hits more
  than a handful of calls on the same flow in a short window — treat
  it as a provider outage and flip the flow to DTMF first.
* The [status page](https://status.orbit.devotel.io) shows a provider
  incident — quote the `request_id` anyway; it still pinpoints the
  affected calls.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.