> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Batch Safety — concurrency-tuner for SMS

> Size bulk SMS batches inside the 10K-per-call platform cap, pick a recipient override when you actually need one, respond to 429 throttling without a retry storm, and re-balance a batch when a sender route fails.

# Batch Safety — the concurrency-tuner for SMS

Sending a bulk SMS is one `POST /api/v1/messages/batch` call with up to 10,000 recipients. What makes it a *safe* bulk send is knowing what the platform gate does for you, where the recipient cap sits tonight, and how to react when a 429 response or a failed sender-route says "slow down" instead of doubling down. This guide is the operational runbook for that surface — pick the batch window, respect the backpressure, and fail over gracefully.

For the supporting terminology on rate limits, see [Rate limits](/guides/rate-limits). For the batch request contract with per-row `variables` and `metadata`, see the [Messages API](/api-reference/endpoints/messaging) reference. For fallback chains when the primary sender route cannot deliver, see [Fallback chains](/guides/fallback-chains).

## 1. What the concurrency-tuner is protecting

Every recipient in a `POST /api/v1/messages/batch` call is first persisted as a `pending` row in the tenant schema, then the send pipeline moves it to `sent`, `scheduled`, or `failed`. The pipeline is **sequential by design** — concurrency past the carrier-allocated SMPP (Short Message Peer-to-Peer) TPS gate would only queue on the upstream provider edge and add throttling noise, and parallel first-hop sends would collide in the tenant's per-second usage counter. The batch surface's safety net therefore has three parts:

* **A recipient-count gate** — rejects calls above the effective effective cap with 422 before any row is inserted.
* **A pre-flight cost estimate** — checks that the org's credit covers the batch's floor cost, returning 402 if not, so an underfunded batch stops at admission and not mid-send.
* **Persist-then-attempt ordering** — on a per-recipient failure, the row flips to `failed` with a readable message rather than leaving a ghost `pending` row.

The concurrency-tuner you can adjust is the **recipient cap** — not the sequencing. The sequencing is load-bearing for everyone.

## 2. Choosing a batch window

The platform default is **10,000 recipients per batch call** — sized so the worker pool stays predictable and the batch API stays cheap as a front-door to large blasts. Between the 10K default and an absolute safety ceiling of 1,000,000 recipients, the platform enforces the clamp in the service — a misconfigured override value can raise the cap but never lower it below the 10K legal floor, and never exceed the 1M ceiling.

Choose the effective cap intentionally:

* **Up to 10K per call** → the platform default. No configuration needed.
* **10K–1M per call (opt-in)** → set the override in your organization settings: flip the batch-cap opt-in flag on, then set a positive integer recipient limit. The platform tolerates middle values; junk values fall back to the 10K default. See the recipe in [Section 5](#5-recipe-tuning-the-over-the-default-cap).
* **Bigger than 1M** → use the campaigns module. It paginates across multiple worker slots and respects the carrier's allocated per-tenant throughput instead of racing one big batch through the API pod.

A per-call window is an upper bound, not a target. Set the override only when a single batch call genuinely needs more than 10K recipients; a middle value you set one Tuesday has consequences for the next burst window.

## 3. Throttling signals: when the provider says slow down

A bulk send hits the same HTTP backpressure contract as any other Orbit call. Respect it instead of fighting it:

| Signal | Where it lands | What it means |
| - | - | - |
| `429 Too Many Requests` | HTTP status + `RATE_LIMITED` code | The API pod's per-key window is exhausted. |
| `Retry-After` (seconds) | 429 header | Wait at least this long before retrying. |
| `X-RateLimit-Limit` / `X-RateLimit-Remaining` / `X-RateLimit-Reset` | every response | Live quota readout — the reset timestamp is when the next full window opens. |
| Partial-batch 207 | HTTP status | At least one recipient failed AND at least one succeeded — the per-recipient outcome map names who. |

The batch endpoint caps parallel upstream pressure for you. A 429 on a batch call is almost always a **client-side over-lapping** signal — the same integration double-firing a cron, a retry loop that ignores `Retry-After`, or a dashboard bulk-action re-run racing a previous call. **Do not** fan the same recipient list into N parallel batch calls to "go faster" — the sequential per-recipient gate already paces your call, and parallel calls multiply the same quota against your key instead of the carrier's. If a 429 appears, the right response is `sleep(retry_after)` and one retry with the same idempotency key, not a loop.

## 4. Re-balance on failover to a secondary sender

A per-recipient `failed` row with a sender-route error is the prompt to re-check routing *before* you retry the row. The batch endpoint does not roll the entire call back because a single upstream gateway hiccuped — it returns the per-recipient outcome map (200 on all-OK, 207 on partial), marks failures on the row, and lets **you** decide whether the next attempt should go through the same sender or a fallback.

* **Read the outcome map before you script the retry.** A 207 means many rows succeeded — re-running the whole batch re-sends them. Instead, re-send only the failed `message_id`s, or use the dashboard's bulk retry on the failed selection (which carries the same 10-at-a-time client concurrency bound).
* **Flip the send route, not the whole batch.** Sender-route failover happens on the routing layer (fallback chains, the smart-send cascade, or a campaign-level channel ladder) — see [Fallback chains](/guides/fallback-chains) for the cascade contract. If the failure class is "primary route is down," re-target the next sender in your chain, and then re-send the failed subset. If the failure class is "recipient rejected" (validation, consent, quiet-hours), fix the recipient data before any route change helps.
* **Use an idempotency key on the retry.** Retrying a failed subset with the same Idempotency-Key replays the prior settled outcomes instead of re-charging the successful ones; a distinct key on a genuinely new attempt remains correct.

## 5. Recipe: tuning the over-the-default cap

The batch recipient limit lives in **two nested keys** under organization settings — both tenant-owned, both default-OFF:

```json theme={null}
{
  "batch": {
    "cap_override_enabled": true,
    "max_recipients": 1000000
  }
}
```

| Field | Value | What it controls |
| - | - | - |
| `cap_override_enabled` | `true` / `false` | The compliance opt-in gate. Until this is explicitly `true`, every batch call is capped at the 10K platform default regardless of what `max_recipients` says. |
| `max_recipients` | integer | The per-call effective cap. Clamped to `[10_000, 1_000_000]`: values below 10K never reduce the floor; values above 1M clamp at the safety ceiling; non-positive / junk values fall back to the platform default. |

With the flag **off**, the `max_recipients` integer is advisory. With the flag **on** and `max_recipients = 1000000`, a single `POST /api/v1/messages/batch` call can carry the 1M-recipient ceiling — sized as a legal bound for one bulk call, past which the right tool is the campaigns module. If a lookup error (DB blip, cache-down) hits while the platform is resolving your override, the default 10K cap protects the worker pool; the override never raises the effective cap above the legal ceiling even on a mis-fetch.

## 6. Comparing against the Message Delivery Report

Sending and *getting through* are two different checkpoints; the batch endpoint's 200/207 verdict reports admission success, and the per-recipient outcome map names which rows were `sent`, `scheduled`, or `failed`. After a batch goes out, the real delivery rate lives in the **Insights → Reports → Message Delivery Report** grid (per channel, per campaign, per window). Compare what the batch API admits against what the report delivers — a partial-batch 207 that "goes green" is the first signal to audit sender-route health and recipient quality, before the batch grows. See [Insights overview](/insights/overview) for that grid's layout.

## 7. Runbook: sustained rate-limit alarms

Use this playbook when either the API returns 429 across several batch windows, or the Insights delivery grid shows a sustained dip after a batch call:

1. **Stop retry-fanning.** Cancel any client-side loop that fires the same recipients until you have a backoff schedule — every extra call costs quota without fixing the gateway.
2. **Read `Retry-After` and obey it.** The header tells you when the API window resets. Treat it as the contract, not as a hint.
3. **Count the effective cap.** Confirm the org settings override hasn't silently widened the window past what the queue can drain.
4. **Check the per-recipient outcome map.** A 207 partial-batch means *some* recipients *did* fail — re-aim only those.
5. **Check sender-route health before rebalancing.** Look at the fallback chain's readiness, not only the batch verdict.
6. **Use the campaigns module past 1M.** If the traffic is truly larger than the ceiling, a paginated campaign distributes the load across workers — the single-call batch never will.
7. **Re-batch only when the alarm clears.** Once 429s stop and Insights shows the delivery rate back to baseline, adopt a lower, sustainable batch window.

## See also

* [Rate limits](/guides/rate-limits) — quotas, headers, retry design.
* [Fallback chains](/guides/fallback-chains) — cascade across channels and routes.
* [Campaign end-to-end](/guides/campaign-end-to-end) — paginated big-batch alternative.
* [Insights overview](/insights/overview) — Message Delivery Report and the Reports grid.
* [Search message history](/guides/search-message-history) — find row-level outcomes after a batch.
