Skip to main content

The per-provider DLR webhook fleet and the shared status pipeline

A carrier delivery receipt (DLR) reaches Orbit on one of a fleet of dedicated webhook endpoints — one per provider. What those endpoints share is everything downstream of the wire format: a single pipeline that verifies the event, commits the message-row status, retries through a bounded queue on a transient miss, and runs the wallet refund arm when the carrier reported a terminal failure. If you integrate against delivery status — webhooks, the message status map, or billing reconciliation — the fleet-versus-shared split is the thing that decides where a behaviour lives and whose rules it follows. For where each status can move next, see Delivery lifecycle and Message status transition rules. For how DLRs leave Orbit back out to your bind, see Send-side DLR model. This page covers the inbound side — who hands a receipt in, and the one pipeline every one of them uses.

The fleet: one endpoint per provider

Orbit exposes a dedicated DLR route for every channel that can return a receipt: the SMPP carrier path, plus the per-channel adapters for WhatsApp, email, fax, RCS (dotgo), and the social channels (Viber, Instagram, Messenger, Discord, Telegram, WeChat, Kakao, LINE, Zalo). Each endpoint is small on purpose — it answers one question: what does this provider’s wire format decode to? A provider adapter belongs on the fleet because its receipt shape is genuinely its own. A WhatsApp status event is not shaped like an SMS carrier deliver_sm, and neither one is shaped like a fax completion notice or a Resend bounce. The boundary of the adapter is exactly provider wire format in, canonical event out. Everything past that layer — verification, commit ordering, retry, refund — is deliberately not per-provider. That per-provider/per-fleet split is the design choice. A provider without its own adapter would leak its native vocabulary into every downstream consumer. A pipeline duplicated per provider would fork on commit ordering, retry cadence, and refund ledger rules. The fleet absorbs the first problem; the shared pipeline prevents the second.

The shared pipeline

Every endpoint in the fleet terminates in the same hand-off, run in this order:
  1. Verify and deduplicate. Replayed deliveries of the same receipt drop before anything else runs. A provider re-firing the same event gets the same ack, and the row sees the transition once.
  2. Validate the row and lock it. The message row is read under a row-level lock and Zod-validated against a fixed schema, so a renamed or missing column stops the write with a loud error instead of silently mis-pricing a refund or mis-gating a regression guard.
  3. Commit the status — or retry. If the row cannot commit at that moment (a tenant-schema blip, a Postgres failover mid-write), the event goes onto the internal DLR retry queue rather than being dropped. The provider receives a 2xx either way, so a transient commit failure never echoes back to the carrier as an error.
  4. Resolve in the canonical vocabulary. The provider’s native status decodes into one of the six canonical delivery states — sent, delivered, read, failed{code}, expired, rejected{code} — before the row is written.
  5. Run the refund arm on terminal failure. When the decoded status is failed, undelivered, or rejected, and the message row carried a pre-send charge, the wallet refund runs as part of the same hand-off, guarded below.
  6. Fan out. Once the commit lands, the webhook event is dispatched to your subscribers, and any downstream triggers (multi-channel fallback, retention policy, verify-channel fallback) fire on the committed transition.
Steps 1–4 are the carrier plane at work; step 5 is billing; step 6 is your webhook. They run as one pipeline so no provider adapter re-implements any of them.

The retry queue: bounded backoff, not the provider’s retry

When step 3 parks an event on the internal queue, the consumer re-attempts the same status update on a widening schedule — exponential backoff keyed on attempt number, capped at 30 seconds between tries, with a hard cap of 30 attempts before the event is treated as undeliverable. That queue is ingress-side: it covers the hop into Orbit, before your subscriber webhook was ever in play. Do not confuse this queue with the egress retry queue in Webhook delivery semantics. That one re-delivers an event to your endpoint after the status committed, when your server could not accept it. The two queues sit at opposite ends of the status lifecycle and a receipt can legitimately ride either, or both, on its way through. The provider’s own webhook-retry loop is a third, unrelated thing — the internal queue exists precisely so the provider does not have to retry. The fan-out option to skip the retry queue entirely exists for one caller class: a provider that deliberately fires the same receipt at a fan-out of candidate tenants (the WhatsApp test-account DLR, which tries each verified test org until one matches). For that class the missing-row case is expected on all but one tenant, and enqueuing a 30-attempt retry per non-match would flood the queue for nothing. Every other caller enqueues normally.

The refund arm

The refund on a terminal failure is not a billing afterthought — it is part of the same hand-off, with the guarantees below wired into the pipeline:
  • Idempotency marker. The row’s refunded_at metadata marker short-circuits a second refund; a replayed DLR or a retried queue job cannot re-credit the wallet.
  • Attempt-scoped idempotency key. The refund ledger key derives from the message’s current external_id, so a stale DLR for a prior attempt cannot refund the original charge a second time after the row was already retried.
  • Charge-stamp-aware amount. The credited amount comes from the charge_amount_cents stamp the pre-send wrote, not from the ceiling the rate table would suggest — a sub-cent charge refunds exactly what left the wallet, and a legacy row with no stamp falls back to the message price.
  • One ledger row. The explicit ledger row and the wallet balance write share one derived transaction id, so the audit record and the credit-back cannot diverge into two entries.
The refundable statuses are exactly the terminal failure set — failed, undelivered, rejected. A sent or delivered receipt never reaches the refund arm. A channel that bills voice minutes stamps its refund under a voice_refund ledger type, so finance can split retransmission credits from carrier refunds in reconciliation. When the refund itself fails, the failure is reported on a documented metric rather than silently swallowing the credit — the customer-visible charge stays on the ledger and is scooped up by reconciliation, not lost.

What this model pins for the reader

Three rules follow from the fleet-versus-shared split, and the naming tells you which one a given behaviour belongs to:
  • A receipt that arrives on the wrong provider’s endpoint gets nowhere. The adapter layer is per-provider by construction; the pipeline only ever sees a canonical event.
  • An adapter that skips the shared pipeline is a defect, not a feature. Status commit, retry, and refund are not re-implementable per provider; a fleet endpoint that ran a second copy of them would fork the commit ordering and the refund ledger rules.
  • The provider’s retry contract never substitutes for the internal queue. The provider is acked the moment the event is either committed or parked; the queue exists so a transient commit failure does not echo back to the carrier.