Skip to main content

Set up a voice broadcast with SMS fallback and per-recipient delivery tracking

This walkthrough plans the whole voice → SMS campaign from the dashboard side: when voice is worth it, which audience source to use, how to attach a recorded clip or a TTS script, how the voice → SMS fallback chain wires up, and how to read every recipient’s delivery outcome. The broadcast composer lives under Outbound → Voice Broadcasts; the fallback hop is a /notify waterfall binding — the clearest shape for “call first, text the rest.” Use this page for an end-to-end campaign plan that starts in the dashboard. For the API-side per-hop receipt, hop-by-hop error codes, and the Inbox reply loop, see Voice broadcast with SMS fallback, threaded back into the Inbox — this page is its setup companion. For a plain one-shot broadcast with no fallback, stay on the voice broadcasts guide.

1. When the voice broadcast is worth it — voice vs. plain SMS

Send a voice broadcast when the message needs a spoken tone, when the wrong channel loses the traffic, or when a keypress (DTMF) wants to feed an IVR branch. The surface is Outbound → Voice Broadcasts — every broadcast it creates is a standard channel: "voice" campaign and passes through the same approvals, wallet, and TCPA gates the full campaign wizard enforces. Plain SMS wins when the message is short, static, and the recipient answers on text anyway. The Outbound → Direct Send flow is the one-shot equivalent: pick a number, pick a template, send now — the text-side shortcut when no listener ever needs to hear the message out loud. The voice → SMS fallback chain is the working middle: voice leads, and when the call cannot land (no-answer, busy, a carrier-level failure), the recipient hears from you by SMS instead. Plan the voice leg first — it is the expensive hop — and let the SMS leg absorb the undelivered cohort. The rest of the page covers that pattern only; if a campaign has real fallback legs on three or more channels, see the fallback channels recipe instead.

2. Pick the recipient source — list, segment, or one-shot CSV

The campaign’s recipient cohort comes from one of three sources; the composer lets you pick any of them in the Recipients control. Build and manage them under Outbound → Audiences before you open the composer.
  • A CDP segment. The living choice: pull a segment the CDP computes from your contact and event data, and the cohort refreshes on every send. Build the segment under Audience → Segments first; the broadcast composer previews its eligible voice-reachable count at send.
  • A contact list. A saved list (static or refreshed manually) built in Audience → Lists. Use it when the cohort is a specific, stable set of recipients — a customer launch cohort, an event reminder list — not a computed property.
  • A one-shot CSV or import. Upload a CSV under Audience → Import for one-off outreach; the import pipeline resolves fields and deduplicates against existing contacts before the send runs.
The composer runs the audience preview (POST /campaigns/audience/preview with channel: "voice") against whichever source you pick. It reports the eligible recipient count and the excluded cohorts — no phone, opted out of calls, suppressed — so a zero-eligible audience never burns a send. Scrub DNC against your tenant do-not-call source before the first send of the day; excluded recipients surface as skips, not silently dropped calls.

3. Attach the message — recorded audio, TTS script, or a voice clone

The composer ships three audio-source paths; the same three load under the dashboard’s media surfaces.
  • TTS script. Type the spoken script into the composer; {{first_name}} and other contact tokens resolve per recipient. The TTS voice picker is the curated stock-voice list plus any voice clone the account holds. To reach the same template from the messaging side (for reuse on the SMS hop), the Messages → Batch composer and the Messages → Templates library carry the same contact tokens.
  • Recorded audio. Upload an MP3, WAV, or M4A clip up to 25 MB. The clip lands in the voice asset store and the resulting handle plays at dispatch; the composer’s Generate preview renders it so you listen back before the send. Browse and manage the same uploaded clips from the Media Library surface (media library guide).
  • Voice clone. Speak the script in a cloned voice — create the clone in Voice → Voice Clones (signed consent and a short sample script), then pick the clone id from the composer’s TTS voice picker. The voice clones guide covers the consent and quality gates; the broadcast composer only references the clone id.
For the SMS fallback hop, the SMS script is not a text-rendering of the voice script — restate the purpose in SMS terms and append opt-out phrasing (Reply STOP to opt out). The SMS hop may be the recipient’s only touch with the message; the SMS template must stand alone.

4. The fallback chain — voice leads, SMS catches the undelivered

The fallback pattern is a /notify waterfall: the voice call is hop 0 (fires immediately), the SMS is hop 1 (armed; fires only when the voice hop reports a terminal failure inside the fallback window). Chain the two bindings like this:
What the request does:
  • mode: "waterfall" — hop 0 (voice) queues immediately; hop 1 (SMS) arms as the escalation tail and fires only when voice terminally fails inside the window.
  • fallback_window_seconds: 600 — the staleness bound. A no-answer, busy, or failed voice receipt inside ten minutes fires the SMS; one that returns eleven minutes later does not escalate.
  • Per-binding max_price — per-hop cost ceiling. A hop whose resolved cost exceeds the cap rejects before the wallet is billed.
  • max_total_price: 0.27 — cumulative cascade cap. A running total that would exceed it trims the escalation tail and reports dropped_for_cost_cap.
  • metadata.campaign — stamped onto every hop’s message row, so cost and delivery reporting attribute both hops to the one campaign (not two orphaned message ids).
A synchronous rejection on the voice hop (unprovisioned sender, an opt-out, a bad from-number) advances the chain in-band before any DLR exists — the POST response names the skipped binding. The SMS hop is the fallback for recipients the call could not reach, not a second channel sent in parallel; mode: "fanout" (the default) fires both bindings at once and is the wrong shape here. The full hop-by-hop receipt reading is in the voice broadcast with SMS fallback walkthrough.

5. Pace, quiet hours, and TCPA — the send-time guardrails

The voice → SMS fallback passes the same send-time gates as any outbound campaign. Verify them before the launch, not after the first 422.
  • TCPA federal window. Every voice hop stamps against the federal 8 AM–9 PM recipient-local dialing window at dispatch; a recipient outside the window refuses with 422 TCPA_FEDERAL_DIALING_WINDOW_BLOCKED. No workspace setting relaxes this. The SMS hop is unaffected and still catches the recipient.
  • Workspace quiet hours. On top of the federal window, your quiet-hours configuration bounds the per-recipient dialing window and is checked at send. The campaign limits and quiet hours guide covers configuring and verifying them; the SMS hop has its own quiet-hours posture (TCPA posture guide).
  • Adaptive pacing. The execution worker throttles per-campaign pacing; the adaptive pacing guide covers the knobs and when pacing relaxes versus tightens. Plan the launch around the recipient cohort’s own local windows, not a single workspace clock.
  • Frequency caps. Workspace-level caps bound how often the same contact receives any outbound send; the composer’s preview excludes capped recipients.

6. Read delivery end-to-end — DLR webhooks and the delivery log

Delivery reporting spans both hops. Wire webhooks before the launch and read the per-recipient outcome across two surfaces.
  • DLR webhooks. Subscribe your endpoint to delivery-report events; the wire DLR webhooks per channel guide carries the per-channel event names and sample payloads, and the DLR webhook consumer guide covers ingestion, signature verification, and retry semantics. The voice hop emits its own receipt (answered / no-answer / busy / failed / AMD-tagged); the SMS hop emits the carrier delivery report when it fires.
  • The Messages delivery log. Per-recipient delivery reads in Messages → Delivery Log — one row per message id, with the message’s channel, status, timestamps, and error code. The delivery log guide reads every column; the developer-side equivalent is developer delivery logs.
  • The campaign rollup. GET /campaigns/:id/stats aggregates the broadcast — answered, failed, delivered count — as DLRs land, and GET /campaigns/:id/voice-cost-rollup reports settled per-minute cost once the broadcast finishes. The campaign ROAS attribution guide interprets the attribution fields.
  • GET /notify/:notifyId. For a single fallback chain, the cascade receipt aggregates hop-by-hop status, per-hop error codes, cost, and which channel landed; sample payloads live in the voice broadcast with SMS fallback walkthrough §4.
The unifying view of both hops, per recipient, is the notify cascade: voice hop recorded, SMS hop recorded, and the recipient’s answered-channel named. The delivery log one row per message; the cascade one envelope per recipient.

7. Compliance — tenant-owned gates before launch

Compliance gates on a voice → SMS fallback are the same ones that gate any outbound campaign; all are tenant-owned (your setting, not a platform-wide switch).
  • 10DLC registration. US SMS traffic to 10DLC numbers rides a registered brand and campaign; the 10DLC registration guide covers the brand + campaign linking. The SMS fallback hop refuses to dispatch on unregistered-to-numbers the same way any SMS send does.
  • TCPA and quiet hours. The federal 8 AM–9 PM recipient-local voice window is non-relaxable; workspace quiet hours are your dialable-window addition on top. The TCPA posture guide walks both the voice federal guard and the SMS-side quiet-hours model.
  • DNC pre-flight scrub. Run the tenant do-not-call source against the recipient list before launching; the DNC preflight scrub guide covers single-recipient and batch scale.
  • Opt-out footer on the SMS hop. The SMS script must include opt-out phrasing (Reply STOP to opt out); the recipient who replies STOP writes to the tenant suppression list and opts out of future sends. The opt-out rules guide covers propagation semantics per channel.
  • Recording-consent acknowledgement. Voice campaigns enforce the recording-consent acknowledgement unconditionally at create time; verify the workspace’s recording-consent posture under the compliance settings first.
  • Approvals. Broadcasts pass the workspace approvals queue when the org requires supervisor sign-off; the outbound approvals guide covers the queue mechanics.
Run this section before the first send of the day, every day: DNC scrub, quiet-hours verification, TCPA window check on the scheduled start.

8. Troubleshooting

  • The SMS never fires. The voice hop answered, or it queued and the terminal failure landed after fallback_window_seconds. Check GET /notify/:notifyId for hop 0’s status and the SMS hop’s absence — a delivered voice hop terminates the chain.
  • 422 TCPA_FEDERAL_DIALING_WINDOW_BLOCKED on part of the cohort. Their local time fell outside the federal window at dispatch. Schedule per recipient-local time, or let the blocked recipients land on the SMS hop (unaffected) — TCPA posture guide.
  • RECIPIENT_OPTED_OUT on the SMS hop. The recipient opted out between the voice failure and the escalation firing, or you did not scrub DNC pre-flight. Verify the tenant DNC source before the next batch.
  • dropped_for_cost_cap names a channel you wanted. max_total_price was lower than the cascade’s resolved cost; raise the cumulative cap.
  • The delivery log shows one row but you expected two. A delivered voice hop terminates the chain — the SMS hop never queued, so there is no second row. This is the waterfall working as designed.

See also