Skip to main content

Canary rollout & shadow dispatch

Promoting a new prompt or model straight to 100% of production traffic is the riskiest way to ship an AI agent change. Orbit gives you two graduated paths instead, and they compose: shadow dispatch proves a candidate on live traffic with zero exposure, and a canary rollout then widens exposure stage by stage while a quality gate watches for regressions. This page explains both paths, the gates they run through, and how to choose between them — the agent versions page owns the version objects these paths act on.

Canary lifecycle: create, split, advance, terminate

A canary rollout takes a saved candidate version and exposes it to a rising share of live traffic, one stage at a time. The rollout is a persisted record — not a query you keep re-supplying — so a scheduler (or an operator clicking “evaluate”) can read where the rollout stands and what it decided. Start one against a candidate at its first stage:
The stages ladder is a list of traffic percents — [5, 25, 50, 100] is the default when you omit it, so you can start with no body beyond the candidate id. Both version ids must be real saved versions of this agent, or the call returns 422. Starting a second rollout while one is active returns 409 ALREADY_ACTIVE — cancel or roll back the active one first. Each evaluation tick loads the candidate’s live quality scorecard — pass rate on outcome rubrics, average sentiment, p50/p95 latency, thumbs-down rate, error rate, computed over a recent window versus an equal-length baseline window — and folds it through the gate: Two sample floors must both clear before a stage widens: the rollout’s own per-stage floor (default 20 conversations, override with min_stage_sample_size at start), plus the scorecard’s own minimum — a detector that declined to judge for lack of data never reads as “clean”. A critical regression bypasses the floors and rolls back immediately; a severe drop is not worth averaging away on a thin window. Warn-level regressions freeze in place by default so the next tick can tell noise from a real trend; set rollback_on_warn: true at start if you want them to revert immediately. Evaluate on demand (the scheduler also ticks this for you):
Two operator kill-switches terminate a rollout regardless of the scorecard: POST /agents/:id/canary-rollout/rollback pulls the candidate (“it went live, and got pulled”), and POST /agents/:id/canary-rollout/cancel voids a rollout started by mistake. GET /agents/:id/canary-rollout returns the persisted record — current stage, status, and the decision history — to any authenticated caller, so dashboard views can render rollout state without write access.

Shadow dispatch: prove a candidate with zero exposure

A shadow run never touches live traffic weights at all. Instead, you point one agent at another: the production agent answers its conversation as usual, and — fire-and-forget after each turn — the runtime replays the same user message and history against the shadow agent and logs both responses. The end user only ever sees the production answer; the shadow answer is purely evidence. Wire up a shadow:
Shadow runs execute in sandbox mode with sandbox: true hardcoded — they never bill your LLM wallet, never write long-term memory, and never dispatch outbound side effects. Two guards block nonsense configurations: an agent cannot shadow itself, and a two-agent cycle is rejected, both with 422. Read the comparison panel:
Each row carries the user message plus both responses, with an exact-match flag and a token-overlap agreement percentage — a sort key for triaging, not a semantic verdict; the operator still reads the pair. The summary aggregates the window: exact-match percent, average agreement, and a token/cost economics block showing what the shadow variant would have cost versus production over the same turns. Those cost figures are advisory — computed locally from token counts, not from the wallet ledger — so use them to compare variants, never for reconciliation. Because the log stores raw chat transcripts, both configure and read calls are restricted to owner, admin, and developer roles. When the evidence looks right, promote:
Promotion copies the shadow’s prompt, model, tools, knowledge bases, and safety config onto the production agent, bumps its version, mints a saved version row so the change lands in the version history and diff surfaces, then clears the shadow link — all in one transaction, so a partial promotion cannot leave the two rows disagreeing. The shadow agent itself stays as a draft, ready to wire up for the next round.

Significance computes; the gate enforces

A/B experiments and canary rollouts both end in a promotion decision, but they split the work across two different mechanisms. The experiment significance engine computes statistics. Given the human_handoff_avoided-style conversion metric you picked at experiment start, it runs a two-proportion chi-square test across the per-arm conversion counts and reports a significance state, and the Bayesian posterior probability that one arm is the better of the two. These numbers are advisory — the detail view returns them; nothing enforces them. The promotion and rollout gates enforce. For A/B experiments, POST /agents/:id/experiments/:expId/promote-winner stays operator-driven, and the auto-promotion evaluator layers four guardrails on top of the statistically recommended winner before an unattended promote is cleared: enough samples per arm, a 0.95 posterior that the winner is actually better, a minimum experiment runtime (default 24 hours, so an early-hours novelty effect cannot auto-ship), and a cost gate. Any failing guardrail holds the promotion and reports which one. After a promote, the auto-rollback evaluator keeps watching: once enough post-promotion impressions accrue, a relative regression versus the pre-promotion baseline (default 10%) reverts automatically. For canary rollouts the scorecard gate described above plays the same enforcing role — but on a broad multi-metric quality surface rather than one conversion metric. The default rule of thumb: reaching for “significance” gets you a number to read; reaching for “gate” gets you something the system will act on.

Where canary and shadow fit against campaign and journey holdouts

Campaign holdouts and journey holdouts measure a different question with the same experimental discipline. A campaign holdout (for example smart_send_holdout_pct on the campaign lifecycle) withholds a slice of the audience from a send so you can attribute lift — “did the message cause conversions?” — after the fact. A journey holdout withholds a control group at journey entry; the journey enrollment fan-out page covers who enters a journey and under what limits. Canary rollouts and shadow dispatch answer a different question on a different surface: not “did this campaign cause lift?” but “is this new agent version safe to put in front of everyone?” Fewer moving parts overlap than the names suggest — holdout audiences never touch agent weights, and canary stages never touch campaign audiences. The shared discipline is real though: both split a population, compare a treatment against a control, and refuse to act on under-powered samples.

Worked sample: shadow a new prompt template, then roll it out

End to end, promoting a rewritten system prompt with minimum blast radius:
  1. Save the candidate version. Edit the prompt on a draft agent (or a fork), then save it as a version on the production agent so it is addressable — see agent versions.
  2. Shadow it. Create a shadow agent carrying the new prompt and point it at production with POST /agents/:id/shadow. Let it collect pairs for a few days, then read GET /agents/:id/shadow-comparison?days=7. Watch the agreement percentage, the exact-match rate, and the economics block — a cheaper variant with matching quality is a strong promote candidate.
  3. Gate the push. If the evidence is clean you have two routes: POST /agents/:id/shadow/promote copies the shadow config onto production in one step (with a fresh saved version), or — when you want exposure to grow gradually — start a canary instead.
  4. Roll out by stages. POST /agents/:id/canary-rollout with candidate_version_id at the default [5, 25, 50, 100] ladder. Each evaluate tick advances cleanly, holds on thin data or warn regressions, and rolls back on a critical regression — with every decision appended to the rollout history you can read back over GET.
  5. Measure after the fact. If you ship via a holdout instead — say a journey where only the treatment branch gets the new agent — the A/B prompt experiments guide covers reading lift and significance against the control.
For agents whose failures show up in quality rather than in any one conversion metric, prefer the canary gate over the experiment significance check: the scorecard it composes watches pass rate, sentiment, latency, satisfaction, and errors at once, so a regression in any of them freezes or reverts the rollout.