Agent A/B experiments: assignment, lift attribution, and promotion
An agent A/B experiment answers one question with live production traffic: does variant B — the candidate prompt — beat variant A — the holdout control — on the one conversion metric you picked, and by how much? This page explains the model: how the split routes traffic, how the assignment stays sticky, how the lift-attribution math guards against an early call, and how the promote decision stays operator-gated. For the step-by-step console and API walkthrough, see Agent prompt A/B experiments: lift, holdout, and attribution.Where experiments sit in the agent lifecycle
An agent change follows a ladder of graduated paths, and a live experiment is the last of them:- Pre-deploy harnesses — sandbox dry-runs, batch simulation, red-team runs, persona simulation. Run these from the console tab before a candidate ever sees traffic (see Pick a testing harness: the decision table).
- Shadow dispatch — replay the production turn against a candidate and log both responses, with zero exposure (see Canary rollout & shadow dispatch).
- Canary rollout — widen a candidate’s traffic share stage by stage behind a quality gate.
- Live A/B experiment — split production traffic between two saved versions and measure the conversion lift.
The experiment object: two versions, one metric
An experiment ties two saved agent versions to one agent and one conversion metric:- Variant A (control) — the frozen production prompt, split at
traffic_split_pctpercent of new contacts. - Variant B (candidate) — the challenger prompt; everything variant A does not take.
- Conversion metric — one of
reply_received,goal_completion, orhuman_handoff_avoided. The experiment counts this one event per assignment; pick it before start.
409 EXPERIMENT_ALREADY_ACTIVE. That single-active invariant is what keeps the attribution numbers honest: there is no cross-experiment traffic leakage to confuse the per-arm conversion counts.
Each conversation that joins the experiment resolves against a version snapshot: the version id is stamped onto the assignment row at first touch, so a later promote or branch of the underlying versions does not retroactively change what variant a given assignment saw.
Assignment: deterministic, sticky, adaptive
The variant picker is deterministic, sticky, and stored — and that is precisely what makes the lift numbers safe to act on.- Deterministic at first touch — the pick is a hash of the experiment id and the contact id: every contact lands in the same arm every time, with no RNG. The usable signal for attribution comes from assignment the first time a contact arrives, not from a fresh random draw on each turn.
- Sticky and stored — once the assignment row for a (experiment, contact) pair is persisted, every future conversation on that contact reads the row back. The picker is never re-rolled.
- Adaptive allocation — the split percent recomputes per new assignment from live per-variant outcomes (Thompson probability-matching by default). During warmup — either arm below 30 impressions — it collapses to the configured
traffic_split_pct, so cold-start is a plain fixed split. The adaptive value is always clamped so both arms keep collecting signal, and stickiness is intact because existing assignments are read back from the store, never re-rolled.
Reading the lift-attribution math
The results payload is computed as pure math from per-variant counts —impressions (assigned conversations) and conversions (metric events) per arm.
- Rates per arm —
conversion_rate = conversions / impressions, with a Wald 95% confidence interval clamped to [0, 1] on each arm’s rate. - Lift — the absolute difference
rateB − rateAin percentage points, plus the relative lift(rateB − rateA) / rateA(null when the control rate is zero). - Two-proportion z-test — significance is a pooled-variance z-test at a two-sided α = 0.05 with 80% power; the p-value is the two-sided tail of the standard normal.
- Wald 95% CI on the difference — each arm’s own unpooled variance, so the A→B gap interval is honest about how far apart the two arms really are.
- Minimum-sample guidance — the service computes the per-variant sample size required to call the observed effect at 95% confidence / 80% power, and reports a
min_sample_reachedflag and astate(ok,collecting, plus the guard states below) so you can tell a real winner from noise.
The promote decision: recommendation, never auto-fire
The experiment never flips the live agent row on its own — promotion stays operator-gated:- Promotion recommendation — the route surfaces a
promotion_recommendationderived from the same posterior model as the split. It recommends a winner only once both arms clear the sample floor (default 100 impressions each) and one arm’s posterior P(best) crosses the threshold (default 0.95):b_clear_winner,a_clear_winner,no_clear_winner, orinsufficient_samples. The actual promotion call is always explicit. - Cost gate — the winner on conversion rate might cost more per conversion (a heavier model, more tool calls). The cost-efficiency layer surfaces the per-variant cost per conversion and flags a hold when the winner is the materially pricier arm (default guard: 25% premium). It is advisory — the gate never blocks on absent cost data.
- Auto-eval quality comparison — reusing the rubric-based LLM-judge eval pipeline keyed by version id, the experiment surfaces each arm’s eval average score / pass rate as another advisory signal. A winner on conversion can quietly degrade response quality; this panel catches it.
- Auto-promotion engine — when configured, the auto-promotion guardrail layers on top of the recommendation (samples per arm, posterior confidence, a minimum experiment runtime so a few early hours never call a winner, and the cost gate). Promotion still surfaces to the operator with every failed guardrail reported by id, and a rollback rule watches the post-promotion window for a material regression.
winner, and mints a fresh history row so the audit log monotonically records what changed.
Where the console renders this
Open Agents → your agent → Experiments. The tab carries the surfaces above as one page:- Active experiment card — variant A vs variant B with per-arm conversion rates, delta, and significance state.
- Lift attribution panel — control vs treatment rates with 95% confidence intervals.
- Results panel — the significance banner, the cost per conversion, and the auto-eval quality comparison.
- Batch simulation panel — a pre-deploy gate run, on the same page so a candidate can be gated before it starts an experiment.
Worked example: a prompt-change experiment
You have a support agent whose refund flow is handled by a long prompt, and you want to test a shorter version. Freeze the production prompt as variant A (apv_control_7), save the candidate as variant B (apv_candidate_12), then:
state — watch for the four guard codes above. The promotion_recommendation field tells you the advisory verdict.
collecting, no_conversions, or the recommendation no_clear_winner — the correct move is to keep collecting until both arms reach the minimum sample, or stop and rerun with a stronger candidate. Do not promote on a tie.
See also
- Pick a testing harness: the decision table — pre-deploy gates to run before an experiment.
- Agent prompt A/B experiments: lift, holdout, and attribution — the operator walkthrough.
- Prompt template lifecycle — how prompt versions are created, slotted, and retired.
- Canary rollout & shadow dispatch — the zero-exposure alternative to a live split.
- Agent run lifecycle — the per-run state machine the experiment sits on top of.
- Agents API endpoints — the REST contract for experiments.