Canary rollout & shadow dispatch
Promoting a new prompt or model straight to 100% of production traffic is the riskiest way to ship an AI agent change. Orbit gives you two graduated paths instead, and they compose: shadow dispatch proves a candidate on live traffic with zero exposure, and a canary rollout then widens exposure stage by stage while a quality gate watches for regressions. This page explains both paths, the gates they run through, and how to choose between them — the agent versions page owns the version objects these paths act on.Canary lifecycle: create, split, advance, terminate
A canary rollout takes a saved candidate version and exposes it to a rising share of live traffic, one stage at a time. The rollout is a persisted record — not a query you keep re-supplying — so a scheduler (or an operator clicking “evaluate”) can read where the rollout stands and what it decided. Start one against a candidate at its first stage:stages ladder is a list of traffic percents — [5, 25, 50, 100] is the default when you omit it, so you can start with no body beyond the candidate id. Both version ids must be real saved versions of this agent, or the call returns 422. Starting a second rollout while one is active returns 409 ALREADY_ACTIVE — cancel or roll back the active one first.
Each evaluation tick loads the candidate’s live quality scorecard — pass rate on outcome rubrics, average sentiment, p50/p95 latency, thumbs-down rate, error rate, computed over a recent window versus an equal-length baseline window — and folds it through the gate:
Two sample floors must both clear before a stage widens: the rollout’s own per-stage floor (default 20 conversations, override with
min_stage_sample_size at start), plus the scorecard’s own minimum — a detector that declined to judge for lack of data never reads as “clean”. A critical regression bypasses the floors and rolls back immediately; a severe drop is not worth averaging away on a thin window. Warn-level regressions freeze in place by default so the next tick can tell noise from a real trend; set rollback_on_warn: true at start if you want them to revert immediately.
Evaluate on demand (the scheduler also ticks this for you):
POST /agents/:id/canary-rollout/rollback pulls the candidate (“it went live, and got pulled”), and POST /agents/:id/canary-rollout/cancel voids a rollout started by mistake. GET /agents/:id/canary-rollout returns the persisted record — current stage, status, and the decision history — to any authenticated caller, so dashboard views can render rollout state without write access.
Shadow dispatch: prove a candidate with zero exposure
A shadow run never touches live traffic weights at all. Instead, you point one agent at another: the production agent answers its conversation as usual, and — fire-and-forget after each turn — the runtime replays the same user message and history against the shadow agent and logs both responses. The end user only ever sees the production answer; the shadow answer is purely evidence. Wire up a shadow:sandbox: true hardcoded — they never bill your LLM wallet, never write long-term memory, and never dispatch outbound side effects. Two guards block nonsense configurations: an agent cannot shadow itself, and a two-agent cycle is rejected, both with 422.
Read the comparison panel:
Significance computes; the gate enforces
A/B experiments and canary rollouts both end in a promotion decision, but they split the work across two different mechanisms. The experiment significance engine computes statistics. Given thehuman_handoff_avoided-style conversion metric you picked at experiment start, it runs a two-proportion chi-square test across the per-arm conversion counts and reports a significance state, and the Bayesian posterior probability that one arm is the better of the two. These numbers are advisory — the detail view returns them; nothing enforces them.
The promotion and rollout gates enforce. For A/B experiments, POST /agents/:id/experiments/:expId/promote-winner stays operator-driven, and the auto-promotion evaluator layers four guardrails on top of the statistically recommended winner before an unattended promote is cleared: enough samples per arm, a 0.95 posterior that the winner is actually better, a minimum experiment runtime (default 24 hours, so an early-hours novelty effect cannot auto-ship), and a cost gate. Any failing guardrail holds the promotion and reports which one. After a promote, the auto-rollback evaluator keeps watching: once enough post-promotion impressions accrue, a relative regression versus the pre-promotion baseline (default 10%) reverts automatically. For canary rollouts the scorecard gate described above plays the same enforcing role — but on a broad multi-metric quality surface rather than one conversion metric.
The default rule of thumb: reaching for “significance” gets you a number to read; reaching for “gate” gets you something the system will act on.
Where canary and shadow fit against campaign and journey holdouts
Campaign holdouts and journey holdouts measure a different question with the same experimental discipline. A campaign holdout (for examplesmart_send_holdout_pct on the campaign lifecycle) withholds a slice of the audience from a send so you can attribute lift — “did the message cause conversions?” — after the fact. A journey holdout withholds a control group at journey entry; the journey enrollment fan-out page covers who enters a journey and under what limits.
Canary rollouts and shadow dispatch answer a different question on a different surface: not “did this campaign cause lift?” but “is this new agent version safe to put in front of everyone?” Fewer moving parts overlap than the names suggest — holdout audiences never touch agent weights, and canary stages never touch campaign audiences. The shared discipline is real though: both split a population, compare a treatment against a control, and refuse to act on under-powered samples.
Worked sample: shadow a new prompt template, then roll it out
End to end, promoting a rewritten system prompt with minimum blast radius:- Save the candidate version. Edit the prompt on a draft agent (or a fork), then save it as a version on the production agent so it is addressable — see agent versions.
- Shadow it. Create a shadow agent carrying the new prompt and point it at production with
POST /agents/:id/shadow. Let it collect pairs for a few days, then readGET /agents/:id/shadow-comparison?days=7. Watch the agreement percentage, the exact-match rate, and the economics block — a cheaper variant with matching quality is a strong promote candidate. - Gate the push. If the evidence is clean you have two routes:
POST /agents/:id/shadow/promotecopies the shadow config onto production in one step (with a fresh saved version), or — when you want exposure to grow gradually — start a canary instead. - Roll out by stages.
POST /agents/:id/canary-rolloutwithcandidate_version_idat the default[5, 25, 50, 100]ladder. Each evaluate tick advances cleanly, holds on thin data or warn regressions, and rolls back on a critical regression — with every decision appended to the rollout history you can read back overGET. - Measure after the fact. If you ship via a holdout instead — say a journey where only the treatment branch gets the new agent — the A/B prompt experiments guide covers reading lift and significance against the control.
Cross-links
- Agent versions — the saved-version objects a canary or shadow promote acts on; versions freeze the prompt and tools a rollout stages.
- Pick a testing harness: the decision table — the pre-deploy sandbox harnesses (dry-run, batch-simulation, red-team, persona-simulation) to run before a candidate ever sees live traffic.
- Agent prompt A/B experiments — the conversion-lift experiment console, the significance stats, and operator-gated winner promotion.
- Campaign lifecycle — holdouts and lift attribution on the marketing side.
- Journey enrollment fan-out — who enters a journey, including control groups you evaluate against.