Batch Safety — the concurrency-tuner for SMS
Sending a bulk SMS is onePOST /api/v1/messages/batch call with up to 10,000 recipients. What makes it a safe bulk send is knowing what the platform gate does for you, where the recipient cap sits tonight, and how to react when a 429 response or a failed sender-route says “slow down” instead of doubling down. This guide is the operational runbook for that surface — pick the batch window, respect the backpressure, and fail over gracefully.
For the supporting terminology on rate limits, see Rate limits. For the batch request contract with per-row variables and metadata, see the Messages API reference. For fallback chains when the primary sender route cannot deliver, see Fallback chains.
1. What the concurrency-tuner is protecting
Every recipient in aPOST /api/v1/messages/batch call is first persisted as a pending row in the tenant schema, then the send pipeline moves it to sent, scheduled, or failed. The pipeline is sequential by design — concurrency past the carrier-allocated SMPP (Short Message Peer-to-Peer) TPS gate would only queue on the upstream provider edge and add throttling noise, and parallel first-hop sends would collide in the tenant’s per-second usage counter. The batch surface’s safety net therefore has three parts:
- A recipient-count gate — rejects calls above the effective effective cap with 422 before any row is inserted.
- A pre-flight cost estimate — checks that the org’s credit covers the batch’s floor cost, returning 402 if not, so an underfunded batch stops at admission and not mid-send.
- Persist-then-attempt ordering — on a per-recipient failure, the row flips to
failedwith a readable message rather than leaving a ghostpendingrow.
2. Choosing a batch window
The platform default is 10,000 recipients per batch call — sized so the worker pool stays predictable and the batch API stays cheap as a front-door to large blasts. Between the 10K default and an absolute safety ceiling of 1,000,000 recipients, the platform enforces the clamp in the service — a misconfigured override value can raise the cap but never lower it below the 10K legal floor, and never exceed the 1M ceiling. Choose the effective cap intentionally:- Up to 10K per call → the platform default. No configuration needed.
- 10K–1M per call (opt-in) → set the override in your organization settings: flip the batch-cap opt-in flag on, then set a positive integer recipient limit. The platform tolerates middle values; junk values fall back to the 10K default. See the recipe in Section 5.
- Bigger than 1M → use the campaigns module. It paginates across multiple worker slots and respects the carrier’s allocated per-tenant throughput instead of racing one big batch through the API pod.
3. Throttling signals: when the provider says slow down
A bulk send hits the same HTTP backpressure contract as any other Orbit call. Respect it instead of fighting it:
The batch endpoint caps parallel upstream pressure for you. A 429 on a batch call is almost always a client-side over-lapping signal — the same integration double-firing a cron, a retry loop that ignores
Retry-After, or a dashboard bulk-action re-run racing a previous call. Do not fan the same recipient list into N parallel batch calls to “go faster” — the sequential per-recipient gate already paces your call, and parallel calls multiply the same quota against your key instead of the carrier’s. If a 429 appears, the right response is sleep(retry_after) and one retry with the same idempotency key, not a loop.
4. Re-balance on failover to a secondary sender
A per-recipientfailed row with a sender-route error is the prompt to re-check routing before you retry the row. The batch endpoint does not roll the entire call back because a single upstream gateway hiccuped — it returns the per-recipient outcome map (200 on all-OK, 207 on partial), marks failures on the row, and lets you decide whether the next attempt should go through the same sender or a fallback.
- Read the outcome map before you script the retry. A 207 means many rows succeeded — re-running the whole batch re-sends them. Instead, re-send only the failed
message_ids, or use the dashboard’s bulk retry on the failed selection (which carries the same 10-at-a-time client concurrency bound). - Flip the send route, not the whole batch. Sender-route failover happens on the routing layer (fallback chains, the smart-send cascade, or a campaign-level channel ladder) — see Fallback chains for the cascade contract. If the failure class is “primary route is down,” re-target the next sender in your chain, and then re-send the failed subset. If the failure class is “recipient rejected” (validation, consent, quiet-hours), fix the recipient data before any route change helps.
- Use an idempotency key on the retry. Retrying a failed subset with the same Idempotency-Key replays the prior settled outcomes instead of re-charging the successful ones; a distinct key on a genuinely new attempt remains correct.
5. Recipe: tuning the over-the-default cap
The batch recipient limit lives in two nested keys under organization settings — both tenant-owned, both default-OFF:
With the flag off, the
max_recipients integer is advisory. With the flag on and max_recipients = 1000000, a single POST /api/v1/messages/batch call can carry the 1M-recipient ceiling — sized as a legal bound for one bulk call, past which the right tool is the campaigns module. If a lookup error (DB blip, cache-down) hits while the platform is resolving your override, the default 10K cap protects the worker pool; the override never raises the effective cap above the legal ceiling even on a mis-fetch.
6. Comparing against the Message Delivery Report
Sending and getting through are two different checkpoints; the batch endpoint’s 200/207 verdict reports admission success, and the per-recipient outcome map names which rows weresent, scheduled, or failed. After a batch goes out, the real delivery rate lives in the Insights → Reports → Message Delivery Report grid (per channel, per campaign, per window). Compare what the batch API admits against what the report delivers — a partial-batch 207 that “goes green” is the first signal to audit sender-route health and recipient quality, before the batch grows. See Insights overview for that grid’s layout.
7. Runbook: sustained rate-limit alarms
Use this playbook when either the API returns 429 across several batch windows, or the Insights delivery grid shows a sustained dip after a batch call:- Stop retry-fanning. Cancel any client-side loop that fires the same recipients until you have a backoff schedule — every extra call costs quota without fixing the gateway.
- Read
Retry-Afterand obey it. The header tells you when the API window resets. Treat it as the contract, not as a hint. - Count the effective cap. Confirm the org settings override hasn’t silently widened the window past what the queue can drain.
- Check the per-recipient outcome map. A 207 partial-batch means some recipients did fail — re-aim only those.
- Check sender-route health before rebalancing. Look at the fallback chain’s readiness, not only the batch verdict.
- Use the campaigns module past 1M. If the traffic is truly larger than the ceiling, a paginated campaign distributes the load across workers — the single-call batch never will.
- Re-batch only when the alarm clears. Once 429s stop and Insights shows the delivery rate back to baseline, adopt a lower, sustainable batch window.
See also
- Rate limits — quotas, headers, retry design.
- Fallback chains — cascade across channels and routes.
- Campaign end-to-end — paginated big-batch alternative.
- Insights overview — Message Delivery Report and the Reports grid.
- Search message history — find row-level outcomes after a batch.