> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Run a live-persona simulation against an agent

> Drive one or more LLM-simulated caller personas turn-by-turn against the agent's sandboxed runtime, then have an LLM judge grade each finished conversation against the scenario's objective and rubric. Returns a rollout-gate report (pass rate, mean score, per-scenario verdicts, tool usage) suitable for gating a prompt/config change in CI. A scenario may instead carry `scripted_utterances` — a deterministic one-utterance-per-turn replay — so the same regression suite re-runs the exact same conversation on every prompt change while the judge still scores it. Use GET .../batch-simulation for fully scripted journeys. Owner/admin/developer role required; 503 when the LLM judge is not configured on the deployment; 404 when the agent does not exist.



## OpenAPI

````yaml /openapi.yaml post /api/v1/agents/{id}/persona-simulation
openapi: 3.1.0
info:
  title: Devotel CPaaS API
  description: Orbit by Devotel — Communications Platform as a Service API
  version: 1.0.0
  contact:
    name: Devotel
    url: https://devotel.io
    email: support@devotel.io
  license:
    name: Proprietary
servers:
  - url: https://api.orbit.devotel.io
    description: Production
security:
  - Bearer: []
  - ApiKey: []
tags:
  - name: Messages
    description: >-
      Send and manage messages across all channels (SMS, WhatsApp, RCS, Email,
      Viber, etc.)
  - name: Agents
    description: AI agent creation, configuration, and execution
  - name: Voice
    description: Voice calls, IVR, conferencing, and SIP trunking
  - name: Webhooks
    description: Webhook endpoint management and delivery logs
  - name: Numbers
    description: Phone number search, provisioning, and configuration
  - name: Contacts
    description: Contact management, segmentation, and lifecycle tracking
  - name: Campaigns
    description: Marketing campaign orchestration and analytics
  - name: Flows
    description: Automation flow builder and execution engine
  - name: Templates
    description: Message template management and approval workflows
  - name: Settings
    description: Organization, channel, and user preference settings
  - name: Verify
    description: OTP generation and verification across channels
  - name: Push
    description: Push notification delivery via FCM and APNs
  - name: Telegram
    description: >-
      Telegram bring-your-own-bot channel: connect a bot, verify a chat against
      the shared sandbox bot, and read unified channel state. Production sends
      go through the Messaging API.
  - name: USSD
    description: >-
      Menu-driven USSD for feature phones: define a menu tree, simulate session
      steps, and host the tenant-scoped aggregator callback.
  - name: Integrations
    description: Third-party service connections and OAuth management
  - name: Files
    description: >-
      Server-to-server media upload, listing, retrieval, and deletion
      (signed-URL backed)
  - name: Messaging Services
    description: >-
      Twilio MessagingService-parity containers bundling sender pool, opt-out
      list, inbound webhook, and sticky-sender/geomatch flags
  - name: Sender Pools
    description: >-
      Group sending numbers into pools with a selection strategy (sticky /
      round-robin / random) for outbound sends
  - name: Opt-Out Lists
    description: >-
      Per-list STOP / HELP / START keyword sets and auto-response copy (Twilio
      Advanced Opt-Out parity)
  - name: SMPP
    description: Tenant BYO-SMPP bind credentials and upstream termination carriers
  - name: Commerce
    description: >-
      Unified omnichannel conversational-commerce: cart/checkout state machine,
      cross-channel payment reconciliation, hosted pay-by-link, native WhatsApp
      checkout, and agentic-checkout payment mandates
  - name: Orby
    description: >-
      In-dashboard Orby operator assistant: streamed assistant turns,
      conversation threads, product knowledge-base search, and the tool-action
      approval gate. Available to signed-in operators only (dashboard session
      auth — API-key requests are rejected).
paths:
  /api/v1/agents/{id}/persona-simulation:
    post:
      tags:
        - Agents
      summary: Run a live-persona simulation against an agent
      description: >-
        Drive one or more LLM-simulated caller personas turn-by-turn against the
        agent's sandboxed runtime, then have an LLM judge grade each finished
        conversation against the scenario's objective and rubric. Returns a
        rollout-gate report (pass rate, mean score, per-scenario verdicts, tool
        usage) suitable for gating a prompt/config change in CI. A scenario may
        instead carry `scripted_utterances` — a deterministic
        one-utterance-per-turn replay — so the same regression suite re-runs the
        exact same conversation on every prompt change while the judge still
        scores it. Use GET .../batch-simulation for fully scripted journeys.
        Owner/admin/developer role required; 503 when the LLM judge is not
        configured on the deployment; 404 when the agent does not exist.
      parameters:
        - schema:
            type: string
          in: path
          name: id
          required: true
        - name: Idempotency-Key
          in: header
          required: false
          description: >-
            Stripe-style idempotency token. Pass a stable, client-generated
            value (1-255 chars) to dedupe retries on transient timeouts. The
            same key+credential+path replays the original response for 24h on
            2xx (5min on 4xx, 30s on 5xx). Returns 409 if a concurrent request
            with the same key is already in flight; replayed responses include
            the `Idempotency-Replay: true` response header.
          schema:
            type: string
            minLength: 1
            maxLength: 255
        - name: X-Test-Mode
          in: header
          required: false
          description: >-
            Sandbox opt-in for Clerk-session-authenticated requests. Set to
            `true` to route the call through the test-mode pipeline: no real
            provider delivery, no credits deducted, response `meta.test_mode:
            true`. **Ignored for live API keys (`dv_live_sk_*`)** —
            server-to-server clients must use a test-prefixed key
            (`dv_test_sk_*`) to exercise sandbox. Test-prefixed keys
            unconditionally enable sandbox regardless of this header.
          schema:
            type: string
            enum:
              - 'true'
              - 'false'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - scenarios
              additionalProperties: true
              properties:
                scenarios:
                  type: array
                  minItems: 1
                  description: >-
                    Simulation scenarios (max 10 per batch). Each carries a live
                    persona — or deterministic `scripted_utterances` — plus an
                    objective and optional rubric the LLM judge grades against.
                  items:
                    type: object
                    required:
                      - id
                      - title
                      - persona
                      - situation
                      - objective
                    additionalProperties: true
                    properties:
                      id:
                        type: string
                      title:
                        type: string
                      channel:
                        type: string
                        enum:
                          - chat
                          - voice
                          - email
                        description: Channel context for the run (default 'chat').
                      difficulty:
                        type: string
                        enum:
                          - easy
                          - medium
                          - hard
                        description: Label for report grouping (default 'medium').
                      persona:
                        type: object
                        required:
                          - name
                        additionalProperties: true
                        properties:
                          name:
                            type: string
                          mood:
                            type: string
                            enum:
                              - neutral
                              - happy
                              - confused
                              - anxious
                              - frustrated
                              - angry
                          background:
                            type: string
                          style:
                            type: string
                          language:
                            type: string
                      situation:
                        type: string
                        description: The caller's context the persona plays.
                      objective:
                        type: string
                        description: What a successful conversation must achieve.
                      rubric:
                        type: array
                        items:
                          type: object
                          additionalProperties: true
                        description: Optional grading criteria (max 10).
                      maxTurns:
                        type: integer
                        minimum: 1
                        description: Optional per-scenario turn cap.
                      scripted_utterances:
                        type: array
                        items:
                          type: string
                        description: >-
                          Deterministic regression-eval mode — one utterance per
                          caller turn, replayed exactly; the persona is ignored
                          but the judge still grades the transcript.
                candidate_version_id:
                  type: string
                  description: >-
                    Score a specific agent version instead of the live one — run
                    the gate BEFORE promoting.
                threshold:
                  type: number
                  minimum: 0
                  maximum: 100
                  description: Pass threshold (0-100, default 70).
                concurrency:
                  type: integer
                  minimum: 1
                  description: Max scenarios run in parallel (default 3).
                judge_model:
                  type: string
                  description: >-
                    Optional LLM judge model override from the Claude allowlist
                    — pin it in regression suites so verdict diffs track the
                    agent change, not judge drift.
            example:
              scenarios:
                - id: scn_refund
                  title: Refund request
                  channel: chat
                  persona:
                    name: Dana
                    mood: frustrated
                    background: Charged twice for the same invoice.
                  situation: A customer wants a duplicate charge refunded.
                  objective: Issue a refund or escalate to billing with a clear handoff.
                  rubric:
                    - key: resolution
                      label: Resolution
                      description: The refund was issued or correctly escalated.
                  maxTurns: 8
              threshold: 75
      responses:
        '200':
          description: >-
            The batch report: per-scenario verdicts plus the aggregate pass/fail
            summary.
          content:
            application/json:
              schema:
                type: object
                additionalProperties: true
                description: >-
                  The batch report: per-scenario verdicts plus the aggregate
                  pass/fail summary.
                properties:
                  data:
                    type: object
                    additionalProperties: true
                    properties:
                      agent_id:
                        type: string
                      candidate_version_id:
                        type:
                          - 'null'
                          - string
                      judge_model:
                        type:
                          - 'null'
                          - string
                      replay_modes:
                        type: object
                        additionalProperties: true
                        description: >-
                          How many scenarios replayed a deterministic script vs
                          a live persona.
                      duration_ms:
                        type: integer
                      summary:
                        type: object
                        additionalProperties: true
                        description: >-
                          Aggregate pass rate, mean judge score, costs, and
                          per-scenario verdicts.
                  meta:
                    type: object
                    properties:
                      request_id:
                        type: string
                        description: >-
                          Unique request identifier (also returned in
                          X-Request-Id header)
                      timestamp:
                        type: string
                        format: date-time
                        description: ISO 8601 UTC timestamp of the response
          headers:
            Idempotency-Replay:
              description: >-
                Set to `true` when the response is a cached replay of a prior
                request with the same `Idempotency-Key`. Absent (or `false`) on
                first-write responses.
              schema:
                type: string
        '400':
          $ref: '#/components/responses/StandardError'
        '401':
          $ref: '#/components/responses/StandardError'
        '403':
          $ref: '#/components/responses/StandardError'
        '404':
          $ref: '#/components/responses/StandardError'
        '422':
          $ref: '#/components/responses/StandardError'
        '429':
          $ref: '#/components/responses/RateLimitedResponse'
        '500':
          $ref: '#/components/responses/StandardError'
components:
  responses:
    StandardError:
      description: >-
        Standard error envelope. `error.code` is machine-readable; see the
        [error reference](https://docs.orbit.devotel.io/reference/error-codes)
        for the catalogue.
      content:
        application/json:
          schema:
            type: object
            properties:
              error:
                type: object
                required:
                  - code
                  - message
                  - status
                properties:
                  code:
                    type: string
                    description: Machine-readable error code (e.g. INVALID_PHONE_NUMBER)
                  message:
                    type: string
                    description: Human-readable error description
                  status:
                    type: integer
                    description: HTTP status code
                  details:
                    type: object
                    additionalProperties: true
                    description: Additional context about the error
              meta:
                type: object
                properties:
                  request_id:
                    type: string
                  timestamp:
                    type: string
                    format: date-time
                  docs_url:
                    type: string
                    format: uri
                    description: Link to relevant error documentation
    RateLimitedResponse:
      description: >-
        Rate limit exceeded. Wait `error.retry_after` seconds (or read the
        `Retry-After` header) before retrying. Returned when the request would
        exceed the bucket identified by `X-RateLimit-Bucket`.
      headers:
        X-RateLimit-Limit:
          description: Total request quota for the current window.
          schema:
            type: integer
            minimum: 0
        X-RateLimit-Remaining:
          description: Requests remaining in the current window.
          schema:
            type: integer
            minimum: 0
        X-RateLimit-Reset:
          description: >-
            Absolute unix-epoch-seconds timestamp at which the current window
            resets.
          schema:
            type: integer
            minimum: 0
        X-RateLimit-Bucket:
          description: >-
            Name of the rate-limit bucket the request was bound by (e.g.
            `auth-write`, `money`, `agent-invoke`, or `custom:<max>/<window>`).
            Stripe-style — lets clients see which named cap they hit.
          schema:
            type: string
        RateLimit-Limit:
          description: >-
            draft-ietf-httpapi-ratelimit-headers no-prefix mirror of
            `X-RateLimit-Limit`. Some SDKs read only this form.
          schema:
            type: integer
            minimum: 0
        RateLimit-Remaining:
          description: >-
            draft-ietf-httpapi-ratelimit-headers no-prefix mirror of
            `X-RateLimit-Remaining`.
          schema:
            type: integer
            minimum: 0
        RateLimit-Reset:
          description: >-
            draft-ietf-httpapi-ratelimit-headers no-prefix mirror — delta
            seconds from now until the window resets (NOT epoch).
          schema:
            type: integer
            minimum: 0
        Retry-After:
          description: >-
            RFC 7231 §7.1.3 — number of seconds the client should wait before
            retrying.
          schema:
            type: integer
            minimum: 1
      content:
        application/json:
          schema:
            type: object
            properties:
              error:
                type: object
                required:
                  - code
                  - message
                  - status
                  - retry_after
                properties:
                  code:
                    type: string
                    enum:
                      - RATE_LIMITED
                    description: >-
                      Always `RATE_LIMITED` — the global request-limiter 429.
                      Distinct from `RATE_LIMIT_EXCEEDED`, the live
                      per-recipient/per-tenant/per-resource frequency-cap 429
                      (not retired).
                  message:
                    type: string
                  status:
                    type: integer
                    enum:
                      - 429
                  retry_after:
                    type: integer
                    minimum: 1
                    description: >-
                      Seconds to wait before retrying. Mirrors `Retry-After`
                      header.
              meta:
                type: object
                properties:
                  request_id:
                    type: string
                  timestamp:
                    type: string
                    format: date-time
                  docs_url:
                    type: string
                    format: uri
  securitySchemes:
    Bearer:
      type: http
      scheme: bearer
      bearerFormat: JWT
      description: Dashboard JWT token from Clerk
    ApiKey:
      type: apiKey
      name: X-API-Key
      in: header
      description: Server-to-server API key (dv_live_sk_*)

````