Skip to main content

Agent A/B experiments: assignment, lift attribution, and promotion

An agent A/B experiment answers one question with live production traffic: does variant B — the candidate prompt — beat variant A — the holdout control — on the one conversion metric you picked, and by how much? This page explains the model: how the split routes traffic, how the assignment stays sticky, how the lift-attribution math guards against an early call, and how the promote decision stays operator-gated. For the step-by-step console and API walkthrough, see Agent prompt A/B experiments: lift, holdout, and attribution.

Where experiments sit in the agent lifecycle

An agent change follows a ladder of graduated paths, and a live experiment is the last of them:
  1. Pre-deploy harnesses — sandbox dry-runs, batch simulation, red-team runs, persona simulation. Run these from the console tab before a candidate ever sees traffic (see Pick a testing harness: the decision table).
  2. Shadow dispatch — replay the production turn against a candidate and log both responses, with zero exposure (see Canary rollout & shadow dispatch).
  3. Canary rollout — widen a candidate’s traffic share stage by stage behind a quality gate.
  4. Live A/B experiment — split production traffic between two saved versions and measure the conversion lift.
The experiment is distinct from the simpler options: a version promote flips the agent row onto a saved version with no measurement, and a canary rollout assumes direction and watches for regressions. An experiment holds the comparison open long enough for attribution math to fire. Each conversation the experiment touches still runs the normal agent run lifecycle — the experiment only changes which version’s prompt the run applies.

The experiment object: two versions, one metric

An experiment ties two saved agent versions to one agent and one conversion metric:
  • Variant A (control) — the frozen production prompt, split at traffic_split_pct percent of new contacts.
  • Variant B (candidate) — the challenger prompt; everything variant A does not take.
  • Conversion metric — one of reply_received, goal_completion, or human_handoff_avoided. The experiment counts this one event per assignment; pick it before start.
Only one experiment per agent can be active at a time — the route rejects a second with 409 EXPERIMENT_ALREADY_ACTIVE. That single-active invariant is what keeps the attribution numbers honest: there is no cross-experiment traffic leakage to confuse the per-arm conversion counts. Each conversation that joins the experiment resolves against a version snapshot: the version id is stamped onto the assignment row at first touch, so a later promote or branch of the underlying versions does not retroactively change what variant a given assignment saw.

Assignment: deterministic, sticky, adaptive

The variant picker is deterministic, sticky, and stored — and that is precisely what makes the lift numbers safe to act on.
  • Deterministic at first touch — the pick is a hash of the experiment id and the contact id: every contact lands in the same arm every time, with no RNG. The usable signal for attribution comes from assignment the first time a contact arrives, not from a fresh random draw on each turn.
  • Sticky and stored — once the assignment row for a (experiment, contact) pair is persisted, every future conversation on that contact reads the row back. The picker is never re-rolled.
  • Adaptive allocation — the split percent recomputes per new assignment from live per-variant outcomes (Thompson probability-matching by default). During warmup — either arm below 30 impressions — it collapses to the configured traffic_split_pct, so cold-start is a plain fixed split. The adaptive value is always clamped so both arms keep collecting signal, and stickiness is intact because existing assignments are read back from the store, never re-rolled.
The holdout discipline is the deliberate choice to freeze your production prompt as variant A. Keep variant A as the unchanging control while B competes; that is what makes the lift you measure meaningful. When you promote a winner, the promote copies the version snapshot into a new history row — the holdout prompt lives on as a saved version.

Reading the lift-attribution math

The results payload is computed as pure math from per-variant counts — impressions (assigned conversations) and conversions (metric events) per arm.
  • Rates per arm — conversion_rate = conversions / impressions, with a Wald 95% confidence interval clamped to [0, 1] on each arm’s rate.
  • Lift — the absolute difference rateB − rateA in percentage points, plus the relative lift (rateB − rateA) / rateA (null when the control rate is zero).
  • Two-proportion z-test — significance is a pooled-variance z-test at a two-sided α = 0.05 with 80% power; the p-value is the two-sided tail of the standard normal.
  • Wald 95% CI on the difference — each arm’s own unpooled variance, so the A→B gap interval is honest about how far apart the two arms really are.
  • Minimum-sample guidance — the service computes the per-variant sample size required to call the observed effect at 95% confidence / 80% power, and reports a min_sample_reached flag and a state (ok, collecting, plus the guard states below) so you can tell a real winner from noise.
Four guard reason codes short-circuit the z-test before it could ever emit a NaN, so the comparison never reads “significant” on degenerate data:

The promote decision: recommendation, never auto-fire

The experiment never flips the live agent row on its own — promotion stays operator-gated:
  • Promotion recommendation — the route surfaces a promotion_recommendation derived from the same posterior model as the split. It recommends a winner only once both arms clear the sample floor (default 100 impressions each) and one arm’s posterior P(best) crosses the threshold (default 0.95): b_clear_winner, a_clear_winner, no_clear_winner, or insufficient_samples. The actual promotion call is always explicit.
  • Cost gate — the winner on conversion rate might cost more per conversion (a heavier model, more tool calls). The cost-efficiency layer surfaces the per-variant cost per conversion and flags a hold when the winner is the materially pricier arm (default guard: 25% premium). It is advisory — the gate never blocks on absent cost data.
  • Auto-eval quality comparison — reusing the rubric-based LLM-judge eval pipeline keyed by version id, the experiment surfaces each arm’s eval average score / pass rate as another advisory signal. A winner on conversion can quietly degrade response quality; this panel catches it.
  • Auto-promotion engine — when configured, the auto-promotion guardrail layers on top of the recommendation (samples per arm, posterior confidence, a minimum experiment runtime so a few early hours never call a winner, and the cost gate). Promotion still surfaces to the operator with every failed guardrail reported by id, and a rollback rule watches the post-promotion window for a material regression.
When you click promote, the route copies the winning version snapshot onto the agent row, stamps the experiment’s winner, and mints a fresh history row so the audit log monotonically records what changed.

Where the console renders this

Open Agents → your agent → Experiments. The tab carries the surfaces above as one page:
  • Active experiment card — variant A vs variant B with per-arm conversion rates, delta, and significance state.
  • Lift attribution panel — control vs treatment rates with 95% confidence intervals.
  • Results panel — the significance banner, the cost per conversion, and the auto-eval quality comparison.
  • Batch simulation panel — a pre-deploy gate run, on the same page so a candidate can be gated before it starts an experiment.

Worked example: a prompt-change experiment

You have a support agent whose refund flow is handled by a long prompt, and you want to test a shorter version. Freeze the production prompt as variant A (apv_control_7), save the candidate as variant B (apv_candidate_12), then:
Traffic now splits deterministically — each contact’s first conversation lands in an arm, and every later conversation on that contact reads the same arm back.
The response carries per-variant stats with Wald CIs, the lift, the z-test p-value, and the state — watch for the four guard codes above. The promotion_recommendation field tells you the advisory verdict.
If the experiment declares the result inconclusive — the state shows collecting, no_conversions, or the recommendation no_clear_winner — the correct move is to keep collecting until both arms reach the minimum sample, or stop and rerun with a stronger candidate. Do not promote on a tie.

See also