> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent A/B experiments: assignment, lift attribution, and promotion

> How a live A/B experiment on an agent works — two saved versions, a deterministic per-contact assignment, lift attribution with Wald 95% confidence intervals and minimum-sample guidance, and the operator-controlled promote-winner decision.

# Agent A/B experiments: assignment, lift attribution, and promotion

An agent A/B experiment answers one question with live production traffic: does variant B — the candidate prompt — beat variant A — the holdout control — on the one conversion metric you picked, and by how much? This page explains the model: how the split routes traffic, how the assignment stays sticky, how the lift-attribution math guards against an early call, and how the promote decision stays operator-gated. For the step-by-step console and API walkthrough, see [Agent prompt A/B experiments: lift, holdout, and attribution](/guides/agent-experiments-ab-prompts).

## Where experiments sit in the agent lifecycle

An agent change follows a ladder of graduated paths, and a live experiment is the last of them:

1. **Pre-deploy harnesses** — sandbox dry-runs, batch simulation, red-team runs, persona simulation. Run these from the console tab before a candidate ever sees traffic (see [Pick a testing harness: the decision table](/concepts/agent-testing-harness-picker)).
2. **Shadow dispatch** — replay the production turn against a candidate and log both responses, with zero exposure (see [Canary rollout & shadow dispatch](/concepts/canary-and-shadow-rollout-model)).
3. **Canary rollout** — widen a candidate's traffic share stage by stage behind a quality gate.
4. **Live A/B experiment** — split production traffic between two saved versions and measure the conversion lift.

The experiment is distinct from the simpler options: a **version promote** flips the agent row onto a saved version with no measurement, and a **canary rollout** assumes direction and watches for regressions. An experiment holds the comparison open long enough for attribution math to fire. Each conversation the experiment touches still runs the normal [agent run lifecycle](/concepts/agent-run-lifecycle) — the experiment only changes which version's prompt the run applies.

## The experiment object: two versions, one metric

An experiment ties two saved agent versions to one agent and one conversion metric:

* **Variant A (control)** — the frozen production prompt, split at `traffic_split_pct` percent of new contacts.
* **Variant B (candidate)** — the challenger prompt; everything variant A does not take.
* **Conversion metric** — one of `reply_received`, `goal_completion`, or `human_handoff_avoided`. The experiment counts this one event per assignment; pick it before start.

Only one experiment per agent can be active at a time — the route rejects a second with `409 EXPERIMENT_ALREADY_ACTIVE`. That single-active invariant is what keeps the attribution numbers honest: there is no cross-experiment traffic leakage to confuse the per-arm conversion counts.

Each conversation that joins the experiment resolves against a version **snapshot**: the version id is stamped onto the assignment row at first touch, so a later promote or branch of the underlying versions does not retroactively change what variant a given assignment saw.

## Assignment: deterministic, sticky, adaptive

The variant picker is deterministic, sticky, and stored — and that is precisely what makes the lift numbers safe to act on.

* **Deterministic at first touch** — the pick is a hash of the experiment id and the contact id: every contact lands in the same arm every time, with no RNG. The usable signal for attribution comes from assignment *the first time a contact arrives*, not from a fresh random draw on each turn.
* **Sticky and stored** — once the assignment row for a (experiment, contact) pair is persisted, every future conversation on that contact reads the row back. The picker is never re-rolled.
* **Adaptive allocation** — the split percent recomputes per new assignment from live per-variant outcomes (Thompson probability-matching by default). During warmup — either arm below 30 impressions — it collapses to the configured `traffic_split_pct`, so cold-start is a plain fixed split. The adaptive value is always clamped so both arms keep collecting signal, and stickiness is intact because existing assignments are read back from the store, never re-rolled.

The **holdout** discipline is the deliberate choice to freeze your production prompt as variant A. Keep variant A as the unchanging control while B competes; that is what makes the lift you measure meaningful. When you promote a winner, the promote copies the version snapshot into a new history row — the holdout prompt lives on as a saved version.

## Reading the lift-attribution math

The results payload is computed as pure math from per-variant counts — `impressions` (assigned conversations) and `conversions` (metric events) per arm.

* **Rates per arm** — `conversion_rate = conversions / impressions`, with a **Wald 95% confidence interval** clamped to \[0, 1] on each arm's rate.
* **Lift** — the absolute difference `rateB − rateA` in percentage points, plus the **relative lift** `(rateB − rateA) / rateA` (null when the control rate is zero).
* **Two-proportion z-test** — significance is a pooled-variance z-test at a **two-sided α = 0.05** with **80% power**; the p-value is the two-sided tail of the standard normal.
* **Wald 95% CI on the difference** — each arm's own unpooled variance, so the A→B gap interval is honest about how far apart the two arms really are.
* **Minimum-sample guidance** — the service computes the per-variant sample size required to call the observed effect at 95% confidence / 80% power, and reports a `min_sample_reached` flag and a `state` (`ok`, `collecting`, plus the guard states below) so you can tell a real winner from noise.

Four guard reason codes short-circuit the z-test before it could ever emit a NaN, so the comparison never reads "significant" on degenerate data:

| `state` | Means |
| - | - |
| `no_traffic` | Both arms have zero impressions — nothing routed yet. |
| `insufficient_variants` | One arm has zero impressions — both need traffic before the test can run. |
| `no_conversions` | Zero conversions across both arms — keep it running. |
| `collecting` | Traffic and conversions exist, but significance or minimum sample not yet reached. |

## The promote decision: recommendation, never auto-fire

The experiment never flips the live agent row on its own — promotion stays operator-gated:

* **Promotion recommendation** — the route surfaces a `promotion_recommendation` derived from the same posterior model as the split. It recommends a winner only once both arms clear the sample floor (default 100 impressions each) and one arm's posterior P(best) crosses the threshold (default 0.95): `b_clear_winner`, `a_clear_winner`, `no_clear_winner`, or `insufficient_samples`. The actual promotion call is always explicit.
* **Cost gate** — the winner on conversion rate might cost more per conversion (a heavier model, more tool calls). The cost-efficiency layer surfaces the per-variant cost per conversion and flags a hold when the winner is the materially pricier arm (default guard: 25% premium). It is advisory — the gate never blocks on absent cost data.
* **Auto-eval quality comparison** — reusing the rubric-based LLM-judge eval pipeline keyed by version id, the experiment surfaces each arm's eval average score / pass rate as another advisory signal. A winner on conversion can quietly degrade response quality; this panel catches it.
* **Auto-promotion engine** — when configured, the auto-promotion guardrail layers on top of the recommendation (samples per arm, posterior confidence, a minimum experiment runtime so a few early hours never call a winner, and the cost gate). Promotion still surfaces to the operator with every failed guardrail reported by id, and a rollback rule watches the post-promotion window for a material regression.

When you click promote, the route copies the winning version snapshot onto the agent row, stamps the experiment's `winner`, and mints a fresh history row so the audit log monotonically records what changed.

## Where the console renders this

Open **Agents → your agent → Experiments**. The tab carries the surfaces above as one page:

* **Active experiment card** — variant A vs variant B with per-arm conversion rates, delta, and significance state.
* **Lift attribution panel** — control vs treatment rates with 95% confidence intervals.
* **Results panel** — the significance banner, the cost per conversion, and the auto-eval quality comparison.
* **Batch simulation panel** — a pre-deploy gate run, on the same page so a candidate can be gated before it starts an experiment.

## Worked example: a prompt-change experiment

You have a support agent whose refund flow is handled by a long prompt, and you want to test a shorter version. Freeze the production prompt as variant A (`apv_control_7`), save the candidate as variant B (`apv_candidate_12`), then:

```bash theme={null}
# 1. Start the experiment (one active per agent)
curl -X POST https://api.orbit.devotel.io/api/v1/agents/agent_abc123/experiments \
  -H "X-API-Key: dv_live_sk_..." \
  -H "Content-Type: application/json" \
  -d '{
    "name": "refund-flow shorter prompt",
    "variant_a_version_id": "apv_control_7",
    "variant_b_version_id": "apv_candidate_12",
    "traffic_split_pct": 50,
    "conversion_metric": "human_handoff_avoided"
  }'
```

Traffic now splits deterministically — each contact's first conversation lands in an arm, and every later conversation on that contact reads the same arm back.

```bash theme={null}
# 2. Read results and lift as the experiment collects
curl "https://api.orbit.devotel.io/api/v1/agents/agent_abc123/experiments/exp_001" \
  -H "X-API-Key: dv_live_sk_..."
```

The response carries per-variant stats with Wald CIs, the lift, the z-test p-value, and the `state` — watch for the four guard codes above. The `promotion_recommendation` field tells you the advisory verdict.

```bash theme={null}
# 3. End the experiment AND promote the winner — both explicit operator calls
curl -X POST https://api.orbit.devotel.io/api/v1/agents/agent_abc123/experiments/exp_001/promote-winner \
  -H "X-API-Key: dv_live_sk_..." \
  -H "Content-Type: application/json" \
  -d '{ "winner": "b", "note": "short prompt wins on handoff avoided" }'
```

If the experiment declares the result **inconclusive** — the state shows `collecting`, `no_conversions`, or the recommendation `no_clear_winner` — the correct move is to keep collecting until both arms reach the minimum sample, or stop and rerun with a stronger candidate. Do not promote on a tie.

## See also

* [Pick a testing harness: the decision table](/concepts/agent-testing-harness-picker) — pre-deploy gates to run before an experiment.
* [Agent prompt A/B experiments: lift, holdout, and attribution](/guides/agent-experiments-ab-prompts) — the operator walkthrough.
* [Prompt template lifecycle](/concepts/prompt-template-lifecycle) — how prompt versions are created, slotted, and retired.
* [Canary rollout & shadow dispatch](/concepts/canary-and-shadow-rollout-model) — the zero-exposure alternative to a live split.
* [Agent run lifecycle](/concepts/agent-run-lifecycle) — the per-run state machine the experiment sits on top of.
* [Agents API endpoints](/api-reference/endpoints/agents) — the REST contract for experiments.
