Skip to main content

IVR flow model and simulator

An IVR flow answers two questions for every caller: what should they hear, and what should the call do with their input — a menu keypress, a spoken phrase, a digit sequence, or a timeout. You author it as a graph of nodes on the IVR studio canvas; the published flow becomes the executable spec the voice runtime walks on every call. This page covers what an IVR flow actually is — the node/verb/intent vocabulary the platform uses — and the full lifecycle of an authored flow: validate the graph, publish an immutable snapshot, resurrect it from the versions catalog, and regression-test it branch-by-branch in the simulator before a live caller ever hits it. The routing decision that selects the flow in the first place is covered by Inbound voice routing; the queues the flow often hands off to are covered by The ACD queue model.

Nodes, verbs, and intents — the vocabulary

Three words come up constantly when authoring, debugging, and reporting on a flow, and they mean three different things:
  • Node — a vertex in the authored graph. Every node has a type (start, play audio, menu, speech input, DTMF input, transfer, voicemail, queue, ring group, time check, language branch, data dip, and so on) and per-type config in its data payload. Nodes are what the canvas, the versions catalog, and the per-node analytics funnel identify by stable id.
  • Verb — a Jambonz instruction the voice edge executes at runtime. Each flow node maps to one verb category: a play or say (text-to-speech) prompt, a gather (DTMF or speech collection), a dial, a record, an enqueue into a queue, or a hangup. The runtime’s verb builder walks the graph and emits the verb sequence for the current step; Jambonz runs those verbs in order and posts results back to the continuation hooks the emit step attached.
  • Intent — a natural-language bucket the classifier resolves a caller’s spoken phrase into: a name plus example phrases, optionally a DTMF-key fallback. Intents are tenant-owned classifier data, distinct from the graph itself. A speech-input node declares which intents it listens for; a matched intent name becomes the edge handle the walker follows.
The simulator’s emitted-verb trace uses the same verb-category projection, which is what lets a regression test assert “this node plays a prompt, that one gathers digits, the transfer dials” without placing a call.

Entry point and the walk

Every flow has one entry node — a start node, or the first interactive node (a prompt or menu) when a legacy graph predates the start tag. From there the runtime walks the graph in two distinct modes:
  • Server-side pass-through nodes — start, time check, language branch, data dip, HTTP request, send message — advance inline. A time check evaluates your holiday/business-hours schedule now; a language branch resolves the caller’s detected or declared language; a data dip looks the caller’s profile up and picks its branch handle. None of these pause for the caller.
  • Interactive nodes — menu, DTMF input, speech input, dial-by-name, and the offer-callback family — pause. The runtime emits a gather verb with a continuation hook, and the walk resumes when the caller’s digits, transcript, or timeout arrives.
Branch precedence is uniform across the whole walker: the caller’s input tries a named edge handle first (a digit, a language tag, an intent name, open/closed, pass/fail), falls back to a default handle, and finally to the single handle-less outgoing edge. A language branch is the exception: an unmatched language with no default handle dead-ends deliberately rather than falling through to an arbitrary edge. Edge resolution and node-type normalization accept both PascalCase and snake_case type tags, which is what keeps flows authored through the SDK walking identically to canvas-authored ones.

Draft to published: the versioning lifecycle

A flow row always carries a mutable draft definition. Working states are private until published, so the canvas can sit half-finished across sessions without ever seeing a live call.
  1. Validate. Before the draft can be persisted or promoted, the platform checks the structural invariants the runtime depends on: exactly one entry node, at least one reachable exit node, no edges pointing at missing nodes, every node reachable from the entry, the flow inside the platform’s node/edge ceilings, and no cycle that cannot reach any terminal. A cycle that loops back on an earlier node but can always escape to a terminal remains valid — retry prompts (“invalid input, try again”) survive the check — a cycle with no escape path is refused at publish time. Validation never throws; the API returns the failure classes as data, and a failed check comes back as a 422 with the offending nodes named.
  2. Publish. Publishing freezes the current draft into an immutable snapshot and pushes it to the published-definition cache the runtime resolves on every call. The published version is what live traffic sees; further canvas edits only touch the draft until you publish again.
  3. Versions catalog. Every publish (and every revert) appends an entry to the flow’s append-only version history. The catalog is a supervisor/audit surface: you can diff any pair of versions or a version against the current draft — added nodes, removed nodes, changed nodes, and whether the compared payload is identical — and revert the flow back to any prior version. A revert collapses to a new draft publish (never a destructive rewrite of history) and itself lands in the catalog, so the lineage stays auditable.
The whole lifecycle runs through a small public surface under voice: creating and editing flows, publishing, listing and comparing versions, and reverting are all standard API calls, scoped so read-only roles can inspect and preview while only write-scoped roles mutate.

The simulator: drive a scripted conversation without a call

The validator asserts structural soundness; the simulator asserts behavioral routing. It walks a draft or published definition exactly the way the runtime walker does — same edge-handle precedence, same pass-through vs interactive split, same loop guard — consuming caller turns you supply for the interactive pauses. A simulation run carries:
  • turns, an ordered script of caller inputs: DTMF digits, an ASR transcript, a forced intent handle, or a forced branch handle for decisions the simulator cannot compute from the graph alone (biometric pass/fail, directory match/no-match, offer accept/decline).
  • Mocked server-side decisions: business-hours open/closed for time checks, a caller language for language branches, and named branch handles for data dips, queue overflow, and voice-biometric verdicts, keyed by node id.
  • Turn consumption accounting — the result reports how many scripted turns the walk consumed and where the script ran out, so an exhausted script at an input node reads as awaiting_input rather than a generic failure.
Every step in the trace records the reached node, the verb categories emitted there, the matched edge handle, the caller turn consumed, and the next node the walk advanced to. The run ends with one of a closed set of termination reasons — natural completion, hangup, dial transfer, queue answered, callback accepted, or a diagnostic class. The diagnostic classes are the regression signals: empty_flow, malformed_definition, no_start_node, node_missing, loop (a node revisited past the runtime’s visit guard), and max_steps (a defense against pathological graphs). A green simulation over a scripted path means the runtime would route a live call the same way, because the simulator deliberately carries the same walk semantics rather than a simplified model. The simulation is a projection over the flow definition: the dial and enqueue labels in the emitted-verb trace describe what the runtime would emit; nothing phones out, and no provider surface is touched. Common use cases are branch regression tests in CI, previewing a draft before publish, and diff drills across versions.

Speech plumbing: intents, hints, and slots

Speech nodes rely on three companion surfaces, each a first-class API:
  • Intent catalog (/voice/ivr-intents) — the tenant-owned bucket of classifier intents. Create, update, and delete buckets, and drive an ad-hoc classify probe against a sample utterance straight from the IVR studio. A speech-input node’s declared prompt phrases and its classifier-backed intent catalog share the same matching semantics: normalized Unicode text, exact phrase match first, then the longest in-sentence phrase match, with a DTMF-key fallback per intent.
  • Recognizer hint catalog (/voice/ivr-nlu-hints) — a flattened, deduplicated, normalized dictionary of utterances the recognizer gets as hints when the speech-input node’s verb is built. Hints raise recognition accuracy for product names, account terms, and short IDs without you retraining the speech model.
  • Slot filling (/voice/ivr-slots) — the stateless per-turn entity extractor. After an intent is known, a self-service flow needs to collect the structured fields fulfilment requires — account number, order ID, date, amount — and re-prompt for whatever is still missing. Each call passes the slot spec and the values collected so far, and gets back the updated values, the next prompt, and whether all slots are filled. The caller (the voice gateway or the IVR studio) threads the accumulating state across turns, so no session table is involved.
All three surfaces read on the read scope and write only on the write scope, and all three isolate per tenant like every other catalog in the platform.

Validation contract and the failure classes the simulator surfaces

Validation runs before persistence, at publish, and is re-checked by the simulator’s preconditions. The failure classes both surfaces name: Cap the investigated failures deliberately at the graph layer: a flow that validates structurally but routes wrong is always a simulator assertion, and a flow that validates behaviorally but still misbehaves live is almost always a mocked server-side decision the run did not parameterize (business hours, language, data dip, biometric verdict).

Where the flow hands the call off

Inbound routing selects a destination type per number before media flows. When that destination is an IVR flow, the flow drives the call until one of four terminal shapes ends the IVR leg:
  • Enqueue into queue — the call parks in a named ACD queue; the presence, skills, and capacity machinery from The ACD queue model picks up dispatch from here, and the hold-music, business-hours, overflow, and callback-family behaviors attach. In the simulator, a queue node terminates with queue_answered unless you mock an overflow handle.
  • Dial / transfer — the flow terminates by bridging to a SIP device, ring group, AI voice agent, or PSTN destination. In the simulator these terminate as dialed.
  • Hangup — an explicit end-of-flow node, terminating as hangup.
  • Voicemail / record — the flow parks the caller in a mailbox capture, after which the recording pipeline attaches.
Once the IVR leg terminates, downstream state — the queue’s SLA timers, a transfer dial’s disposition, voicemail recording state — reports through the per-call record and the node’s funnel analytics. The flow’s job ends at the hand-off; analytics continue per node across calls on its own funnel surface.

Worked sample: play → interactive menu → speech → queue

Consider a support line whose draft flow has four canonical nodes:
  1. A play-audio node that greets the caller and does not pause.
  2. A menu node offering “press 1 or say support; press 2 or say sales” with a speech-input-shaped transcript fallback.
  3. A data dip that looks the caller’s tier up and branches vip/default on a mocked handle in simulation.
  4. A queue node that enqueues the caller into the tier-specific support queue, with a mocked overflow edge looping back to the menu.
A simulation scripted with turns [{ speech: "support" }] and branch mocks { dataDipNodeId: "vip" } reports:
A sibling run on turns [{ digits: "2" }] routes the same graph to the sales queue; one on an empty turn list terminates at awaiting_input on the menu node, which is how a regression suite discovers it forgot a turn. The same graph validated structurally first, then simulated behaviorally, reaches publish safely — and every subsequent publish lands in the versions catalog, so a regression red against a new draft can diff against the last known-good snapshot and revert on the spot.

See also