Troubleshooting: IVR NLU classifier and slot-extraction failures
A conversational IVR menu replaces “press 1 for balance” with two model calls per utterance — a classifier that decides which intent the caller meant, and a slot extractor that pulls entities (an account number, a date, a name) out of what they said. Either call can fail with its own 503:IVR_CLASSIFY_LLM_ERROR— the classifier call failed, so the menu could not decide where the utterance routes.IVR_SLOT_EXTRACT_LLM_ERROR— the slot-extraction call failed, so the menu got a route but could not fill the entities the next node reads.
What a classify failure means
A speech-input node declares the intents it listens for; on each utterance the classifier scores the caller’s phrase against your active intent set and resolves a match.IVR_CLASSIFY_LLM_ERROR
means that scoring call never returned — not that the caller said
something unintelligible, and not a low-confidence fallback_reason
from the intent tester. The walker cannot
pick an edge, so the flow has no route for this turn.
The distinction matters: a low-confidence match is a routing-quality
problem you fix in the intent descriptions. A 503-classify failure is
an infrastructure problem — the model leg of the menu — and you treat
it as a provider blip, not a prompt to tune.
What a slot-extraction failure means
Slot extraction runs after classification, when the resolved intent needs entities (the account number the caller spoke, the date they asked for).IVR_SLOT_EXTRACT_LLM_ERROR means the route decision
succeeded but the entity pull failed, so the next node reads an empty
or partial slot object. A flow that branches on a filled slot (for
example a payment-date lookup) sees the same “no data” shape it would
see if the caller stayed silent.
Recover the live call
A 503 here must never hard-fail the caller into a hangup — pick one of these per flow, in order of preference:- Re-ask the caller. The cheapest recovery on a live call is a re-prompt: “Sorry, could you say that again?” for classification, or re-ask for the specific field on slot failure. One or two re-prompt attempts per turn is the right cap; beyond that, move to the next option.
- Fall back to the DTMF digit menu. If the flow carries a
DTMF mirror for the same menu (“press 1 for billing”), take the
digit path after the first NLU 503. Callers keep moving while the
model leg is down. The NLU routing guide
recommends building every conversational menu with this mirror —
a
defaultedge from the speech node to adtmfInputnode or a queue. - Drop the call to a queue. For flows where a wrong turn is worse than a wait (payments, account changes), route the failure edge to an ACD queue with a human rather than guessing on a stale NLU layer.
Where the LLM call is made
Both codes come from the same gateway the rest of Orbit uses: per the glossary, Orbit is Anthropic-only — every LLM request (agent runtime, NLU classifier, slot extractor) routes through a single partner-vetted Anthropic Claude gateway. Tenants do not connect their own model-provider keys, so there is no BYOK setting for you to flip here; when the classifier 503s across every caller on a flow, treat it as a provider-side event, check the status page, and escalate with therequest_id the error envelope carries.
Flip the flow to DTMF-only
When the NLU path is flapping, take the model leg out of the call path instead of serving failures to callers:- In the flow builder, repoint the speech-input node’s edges at the
DTMF mirror node (or swap the speech node for a
dtmfInputnode); keep the same queue/agent destinations on each branch. - Publish the edited graph — inbound attaches the published snapshot, so no inbound-number changes are needed.
- Re-run the flow simulator’s fallback branch to confirm a digit press lands on the intended queue rather than hanging in the speech node (see IVR flow model and simulator).
- Re-enable the NLU path once the provider incident clears: repoint the edges back to the speech node and republish. Your intents keep their definitions either way — deactivate an intent only to drop it out of classification during tuning, not for an outage flip.
When to escalate
Escalate with therequest_id and the flow name when any of these
holds:
IVR_CLASSIFY_LLM_ERRORorIVR_SLOT_EXTRACT_LLM_ERRORhits more than a handful of calls on the same flow in a short window — flip to DTMF first, then escalate.- The 503s survive your re-prompt cap and DTMF fallback on prompts that classified cleanly before.
- The status page shows a provider
incident — quote the
request_idanyway; it pinpoints the affected calls.