Skip to main content

Cross-channel cascade failover policy

A cascade failover policy is a persisted, per-organization ordered channel chain that attaches to ordinary messaging sends. When the primary channel returns a terminal undelivered receipt inside the freshness window, Orbit automatically dispatches the next channel as a new send, then the next, until one delivers or the chain is exhausted. Every leg — primary and fallback — shares one message_group_id, so the whole run reads as a single logical message. This page is about the persisted policy object and the per-send cascade field on the channel send endpoints (POST /api/v1/messages/{sms,rcs,whatsapp,...}). It is separate from the notify composer cascade, which builds per-recipient waterfall chains inside one POST /api/v1/notify call. Use this page when you want failover behavior attached to your existing sendMessage traffic. Every control below is tenant-owned: the channel order, the freshness window, the per-send opt-out, and the body template are choices you make per organization or per send.

When to choose cascade over parallel broadcast

Choose cascade when the goal is “one message, one recipient, best channel wins” — the richest reliable channel should try first, and cheaper fallbacks should fire only on a terminal failure. Choose parallel broadcast when the goal is “one message, many recipients, right now” and every channel binding should fire independently. The subscribe-many broadcast guide covers that shape: one body, N bindings, N parallel delivery attempts, one response envelope. Cascade, by contrast, is sequential: only the active hop fires, and each subsequent hop waits for the previous one to fail.

Endpoint walkthrough

Read or replace the org policy

Read the current policy:
Replace the policy:
The policy object has three fields: The eligible channel set is closed: sms, whatsapp, rcs, viber, telegram, messenger, instagram, line, apple_messages. Voice, fax, email, and push are never cascade-eligible — their billing and consent semantics differ too much for an automatic hop. Validation rejects anything outside the set with a 422 that names each invalid field.

Apply the policy on a single send

When a send omits its own cascade input, the org policy stamps the fallback chain into metadata. The send response returns the primary message id; that id is the first leg of the logical group.
A normal 200 response carries the primary message_id:
That msg_7f3a2b... is the first leg. The message_group_id is stamped at send time but is not returned in the send envelope; read it back through the group endpoint or the Delivery Log.

Override the policy per send

To supply a chain for one send, use the typed cascade field:
To opt a single send out of the org policy:
Precedence is fixed: a legacy metadata.fallback_channels chain wins over the typed cascade field, and the typed field wins over the org policy. Sources are never merged, so your per-send chain is never silently extended by the org default.

Reading cross-leg state

Resolve every leg of one logical send by its group id:
The response lists legs in send order, oldest first:
The primary leg has fallback_of: null; each fallback leg records the message it fell back from and the terminal status that triggered the escalation. A group read is bounded at 100 legs and returns an empty list for an unknown group id rather than an error. In the dashboard, the Delivery Log groups legs primary-first on the same envelope, so operators see one logical message instead of disconnected sends.

Per-leg DLR and error-code attribution

A fallback hop fires only on a terminal DLR inside the freshness window. The statuses that trigger escalation are failed, undelivered, and rejected. When one lands, the engine dispatches the next channel as a brand-new send and stamps two attribution fields on the new leg:
  • fallback_of — the message id of the leg that failed.
  • fallback_reason — a machine-readable string such as primary_rcs_undelivered, primary_whatsapp_failed, or primary_sms_rejected.
The error code from the failing leg does not propagate into the next leg’s status; the next leg starts fresh at queued. It does, however, re-enter the canonical send pipeline, so sender resolution, opt-out checks, and quiet-hours gates apply per hop. An SMS opt-out blocks an SMS hop even if the primary was RCS or WhatsApp.

Cascade vs notify

Two different surfaces solve two different problems: Use the failover policy when your integration already calls the channel send endpoints and you want cascade behavior without changing endpoint URLs. Use notify cascade when you are broadcasting or fail-fording one message to a list of recipients across channels.

Content considerations

The same body must make sense on the primary channel and the terminal fallback. An RCS rich card or WhatsApp template can carry buttons, images, and variables; SMS plain text cannot. Design your content so the fallback is still usable:
  • OTPs and alert codes — put the code in the first sentence in plain text. Rich formatting can surround it, but the code itself must survive a plain-text fallback.
  • Calls to action — write the URL or phone number in full text, not only as a button. SMS fallback loses interactive buttons.
  • Variables — keep variable count low and order stable. If the WhatsApp template accepts four variables and the SMS fallback uses the same four in the same order, the same template parameter object feeds both.
  • Sender identity — the fallback channel resolves its own sender. Do not rely on a rich-channel sender name that does not map to SMS alphanumeric sender limits.
If the primary and fallback bodies must differ materially, send two distinct messages instead of using cascade, or use the per-hop body override available in notify cascade.

Cost tallying

Each leg bills as its own send on its own channel. There is no separate cascade fee; you pay for the primary plus each fallback that actually fires. The logical group endpoint lists every attempted leg, and billing records are joined by message_group_id so you can roll up spend per logical message. For a full cost breakdown, see the Billing overview. The Delivery Log and cascade group endpoint both show per-hop price and currency; the group response does not sum them — sum them client-side or query the billing ledger by message_group_id.

Idempotency and throttling

  • Idempotency — the original send honors the Idempotency-Key header like any other create call. A terminal DLR triggers at most one fallback per message because the engine atomically stamps fallback_attempted_at on the original row before dispatching the next hop.
  • Cancel — set cascade.enabled: false on a send to prevent that specific send from arming a chain, regardless of the org policy.
  • Fail-shorten — update the org policy to a shorter chain or set enabled: false to stop arming new sends. Already-armed cascades continue with the chain they were stamped with at send time.
  • Rate limits — policy reads carry messages:read or messages:write scope at 60 requests per minute; writes carry messages:write plus owner, admin, or developer role at 20 requests per minute.

Worked example: OTP RCS → WhatsApp → SMS

A tenant policy arms every RCS send to fall back through WhatsApp and then SMS. A customer requests a one-time passcode. Send the primary:
The RCS leg fails with a terminal DLR:
The escalation engine fires the WhatsApp hop:
Resolve the whole run by group id:
The group shows two legs: the failed RCS primary and the delivered WhatsApp fallback, both under the same message_group_id. If WhatsApp had also failed, the engine would have continued to SMS with the same group id and a fallback_reason of primary_whatsapp_undelivered.

See also