Skip to main content
POST
Run a live-persona simulation against an agent

Authorizations

Authorization
string
header
required

Dashboard JWT token from Clerk

Headers

Idempotency-Key
string

Stripe-style idempotency token. Pass a stable, client-generated value (1-255 chars) to dedupe retries on transient timeouts. The same key+credential+path replays the original response for 24h on 2xx (5min on 4xx, 30s on 5xx). Returns 409 if a concurrent request with the same key is already in flight; replayed responses include the Idempotency-Replay: true response header.

Required string length: 1 - 255
X-Test-Mode
enum<string>

Sandbox opt-in for Clerk-session-authenticated requests. Set to true to route the call through the test-mode pipeline: no real provider delivery, no credits deducted, response meta.test_mode: true. Ignored for live API keys (dv_live_sk_*) — server-to-server clients must use a test-prefixed key (dv_test_sk_*) to exercise sandbox. Test-prefixed keys unconditionally enable sandbox regardless of this header.

Available options:
true,
false

Path Parameters

id
string
required

Body

application/json
scenarios
object[]
required

Simulation scenarios (max 10 per batch). Each carries a live persona — or deterministic scripted_utterances — plus an objective and optional rubric the LLM judge grades against.

Minimum array length: 1
candidate_version_id
string

Score a specific agent version instead of the live one — run the gate BEFORE promoting.

threshold
number

Pass threshold (0-100, default 70).

Required range: 0 <= x <= 100
concurrency
integer

Max scenarios run in parallel (default 3).

Required range: x >= 1
judge_model
string

Optional LLM judge model override from the Claude allowlist — pin it in regression suites so verdict diffs track the agent change, not judge drift.

Response

The batch report: per-scenario verdicts plus the aggregate pass/fail summary.

The batch report: per-scenario verdicts plus the aggregate pass/fail summary.

data
object
meta
object