Skip to main content

Batch Safety — the concurrency-tuner for SMS

Sending a bulk SMS is one POST /api/v1/messages/batch call with up to 10,000 recipients. What makes it a safe bulk send is knowing what the platform gate does for you, where the recipient cap sits tonight, and how to react when a 429 response or a failed sender-route says “slow down” instead of doubling down. This guide is the operational runbook for that surface — pick the batch window, respect the backpressure, and fail over gracefully. For the supporting terminology on rate limits, see Rate limits. For the batch request contract with per-row variables and metadata, see the Messages API reference. For fallback chains when the primary sender route cannot deliver, see Fallback chains.

1. What the concurrency-tuner is protecting

Every recipient in a POST /api/v1/messages/batch call is first persisted as a pending row in the tenant schema, then the send pipeline moves it to sent, scheduled, or failed. The pipeline is sequential by design — concurrency past the carrier-allocated SMPP (Short Message Peer-to-Peer) TPS gate would only queue on the upstream provider edge and add throttling noise, and parallel first-hop sends would collide in the tenant’s per-second usage counter. The batch surface’s safety net therefore has three parts:
  • A recipient-count gate — rejects calls above the effective effective cap with 422 before any row is inserted.
  • A pre-flight cost estimate — checks that the org’s credit covers the batch’s floor cost, returning 402 if not, so an underfunded batch stops at admission and not mid-send.
  • Persist-then-attempt ordering — on a per-recipient failure, the row flips to failed with a readable message rather than leaving a ghost pending row.
The concurrency-tuner you can adjust is the recipient cap — not the sequencing. The sequencing is load-bearing for everyone.

2. Choosing a batch window

The platform default is 10,000 recipients per batch call — sized so the worker pool stays predictable and the batch API stays cheap as a front-door to large blasts. Between the 10K default and an absolute safety ceiling of 1,000,000 recipients, the platform enforces the clamp in the service — a misconfigured override value can raise the cap but never lower it below the 10K legal floor, and never exceed the 1M ceiling. Choose the effective cap intentionally:
  • Up to 10K per call → the platform default. No configuration needed.
  • 10K–1M per call (opt-in) → set the override in your organization settings: flip the batch-cap opt-in flag on, then set a positive integer recipient limit. The platform tolerates middle values; junk values fall back to the 10K default. See the recipe in Section 5.
  • Bigger than 1M → use the campaigns module. It paginates across multiple worker slots and respects the carrier’s allocated per-tenant throughput instead of racing one big batch through the API pod.
A per-call window is an upper bound, not a target. Set the override only when a single batch call genuinely needs more than 10K recipients; a middle value you set one Tuesday has consequences for the next burst window.

3. Throttling signals: when the provider says slow down

A bulk send hits the same HTTP backpressure contract as any other Orbit call. Respect it instead of fighting it: The batch endpoint caps parallel upstream pressure for you. A 429 on a batch call is almost always a client-side over-lapping signal — the same integration double-firing a cron, a retry loop that ignores Retry-After, or a dashboard bulk-action re-run racing a previous call. Do not fan the same recipient list into N parallel batch calls to “go faster” — the sequential per-recipient gate already paces your call, and parallel calls multiply the same quota against your key instead of the carrier’s. If a 429 appears, the right response is sleep(retry_after) and one retry with the same idempotency key, not a loop.

4. Re-balance on failover to a secondary sender

A per-recipient failed row with a sender-route error is the prompt to re-check routing before you retry the row. The batch endpoint does not roll the entire call back because a single upstream gateway hiccuped — it returns the per-recipient outcome map (200 on all-OK, 207 on partial), marks failures on the row, and lets you decide whether the next attempt should go through the same sender or a fallback.
  • Read the outcome map before you script the retry. A 207 means many rows succeeded — re-running the whole batch re-sends them. Instead, re-send only the failed message_ids, or use the dashboard’s bulk retry on the failed selection (which carries the same 10-at-a-time client concurrency bound).
  • Flip the send route, not the whole batch. Sender-route failover happens on the routing layer (fallback chains, the smart-send cascade, or a campaign-level channel ladder) — see Fallback chains for the cascade contract. If the failure class is “primary route is down,” re-target the next sender in your chain, and then re-send the failed subset. If the failure class is “recipient rejected” (validation, consent, quiet-hours), fix the recipient data before any route change helps.
  • Use an idempotency key on the retry. Retrying a failed subset with the same Idempotency-Key replays the prior settled outcomes instead of re-charging the successful ones; a distinct key on a genuinely new attempt remains correct.

5. Recipe: tuning the over-the-default cap

The batch recipient limit lives in two nested keys under organization settings — both tenant-owned, both default-OFF:
With the flag off, the max_recipients integer is advisory. With the flag on and max_recipients = 1000000, a single POST /api/v1/messages/batch call can carry the 1M-recipient ceiling — sized as a legal bound for one bulk call, past which the right tool is the campaigns module. If a lookup error (DB blip, cache-down) hits while the platform is resolving your override, the default 10K cap protects the worker pool; the override never raises the effective cap above the legal ceiling even on a mis-fetch.

6. Comparing against the Message Delivery Report

Sending and getting through are two different checkpoints; the batch endpoint’s 200/207 verdict reports admission success, and the per-recipient outcome map names which rows were sent, scheduled, or failed. After a batch goes out, the real delivery rate lives in the Insights → Reports → Message Delivery Report grid (per channel, per campaign, per window). Compare what the batch API admits against what the report delivers — a partial-batch 207 that “goes green” is the first signal to audit sender-route health and recipient quality, before the batch grows. See Insights overview for that grid’s layout.

7. Runbook: sustained rate-limit alarms

Use this playbook when either the API returns 429 across several batch windows, or the Insights delivery grid shows a sustained dip after a batch call:
  1. Stop retry-fanning. Cancel any client-side loop that fires the same recipients until you have a backoff schedule — every extra call costs quota without fixing the gateway.
  2. Read Retry-After and obey it. The header tells you when the API window resets. Treat it as the contract, not as a hint.
  3. Count the effective cap. Confirm the org settings override hasn’t silently widened the window past what the queue can drain.
  4. Check the per-recipient outcome map. A 207 partial-batch means some recipients did fail — re-aim only those.
  5. Check sender-route health before rebalancing. Look at the fallback chain’s readiness, not only the batch verdict.
  6. Use the campaigns module past 1M. If the traffic is truly larger than the ceiling, a paginated campaign distributes the load across workers — the single-call batch never will.
  7. Re-batch only when the alarm clears. Once 429s stop and Insights shows the delivery rate back to baseline, adopt a lower, sustainable batch window.

See also