Run a scripted simulation batch with a regression gate
Play a batch of scenarios — each with a fixed caller utterance script — through the agent in sandbox mode (no billing, no memory writes), have the LLM judge grade each finished transcript against the scenario’s objective and success criteria, then run the batch summary through a regression gate that returns a pass/fail signal with reasons. Use it as a CI or pre-promotion quality gate: the deterministic script makes scores comparable across prompt, model, or tool changes, unlike a live-persona simulation. Body: scenarios (id, title, situation, objective, scripted_utterances, optional persona/success criteria/max turns), plus optional candidate_version_id, judge threshold, concurrency, and gate knobs min_pass_rate / max_error_rate. Responds 404 when the agent is unknown and 503 when no LLM judge is configured. Owner, admin, or developer role required.
Authorizations
Dashboard JWT token from Clerk
Headers
Stripe-style idempotency token. Pass a stable, client-generated value (1-255 chars) to dedupe retries on transient timeouts. The same key+credential+path replays the original response for 24h on 2xx (5min on 4xx, 30s on 5xx). Returns 409 if a concurrent request with the same key is already in flight; replayed responses include the Idempotency-Replay: true response header.
1 - 255Sandbox opt-in for Clerk-session-authenticated requests. Set to true to route the call through the test-mode pipeline: no real provider delivery, no credits deducted, response meta.test_mode: true. Ignored for live API keys (dv_live_sk_*) — server-to-server clients must use a test-prefixed key (dv_test_sk_*) to exercise sandbox. Test-prefixed keys unconditionally enable sandbox regardless of this header.
true, false Path Parameters
Response
Successful response.