> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Run your first QA evaluation, from rubric to coaching plan

> Build one rubric, score a recorded call, work its appeal, and follow the result into coaching and practice.

# Run your first QA evaluation, from rubric to coaching plan

Use one recorded support call to connect the parts of a first-week QA setup: define a small rubric, turn on automated scoring and sampling, review a flagged result, let the agent respond, and follow a low score into coaching. The same rubric stays attached to the call from first score through the coaching follow-up.

You need an active voice team, a recorded call with a transcript and an assigned agent, and an owner, admin, or supervisor account for setup and review. The example values below are illustrative; replace IDs with values from your workspace. Auto-score is tenant-controlled and off until you enable it.

## 1. Author one two-section rubric

In the dashboard, open **Quality → Rubrics** and create **Support call – first review**. Add two sections:

* **Required checks**: one identity-verification criterion, set as the section's `auto_fail` compliance gate.
* **Conversation quality**: one criterion for clear next steps.

Set the weights to 1 and 2 respectively. This makes clear next steps count twice as much as identity verification when the call passes the required check. A zero on identity verification forces the evaluation total to zero.

Create the same form over the QA API. The `auto_fail` flag belongs to the compliance section; the criterion remains individually scoreable.

```bash theme={null}
curl -X POST https://api.orbit.devotel.io/api/v1/quality/evaluation-forms \
  -H "X-API-Key: $ORBIT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Support call – first review",
    "sample_rate_pct": 10,
    "definition": {
      "sections": [
        {
          "id": "required-checks",
          "label": "Required checks",
          "weight": 1,
          "auto_fail": true,
          "criteria": [
            { "id": "identity-verified", "label": "Verified identity before account details", "weight": 1 }
          ]
        },
        {
          "id": "conversation-quality",
          "label": "Conversation quality",
          "weight": 2,
          "criteria": [
            { "id": "clear-next-steps", "label": "Explained the next step and owner", "weight": 1 }
          ]
        }
      ]
    }
  }'
```

Save the returned form ID as `FORM_ID`. Keep this as your active form for the walkthrough: AI Auto-QA uses the newest active evaluation form. The 10% form sample rate is a starting point for a small review sample, not a promise that exactly one of ten calls will appear in every short window.

For the full weighting, calibration, and form lifecycle guidance, see [Build a contact-center QA program](/guides/quality-management-program).

## 2. Enable auto-score and sampling

Open **Settings → QA auto-score**, enable the tenant setting, and set the flag threshold to **70**. Then open **Settings → QA sampling** and start with a modest **10%** sample. Auto-score evaluates eligible completed calls; sampling provides a separate human-review sample. They work together, but a sampled unscored row and an AI-scored flagged row are different queue items.

The first call must be completed, have a transcript, and have an assigned agent. After the next autoscore pass, open **Quality → Evaluations** and filter for auto-scored or flagged evaluations. Confirm that one row points to the recorded call and shows `auto_scored: true`. If no row appears, check the auto-score toggle, active form, transcript, agent assignment, and [QA auto-score settings](/guides/qa-autoscore-settings). For how the sample share is configured, see [QA sampling settings](/guides/qa-sampling-settings).

For this worked example, the scorer rates **identity verified = 100** and **clear next steps = 40**, for a weighted total of about **60/100**. Because 60 is below the 70 flag threshold, the evaluation is marked for a human review. The auto-score is a first pass; the reviewer owns the final human assessment.

## 3. Review and grade the flagged call

In **Quality → Evaluations**, select the flagged row for the example call. The reviewer view brings together the call and transcript, the rubric sections and criteria, the automated scores, and the review flag. Listen to the relevant part of the recording before changing a score. In this example, the next step was unclear, so keep identity verification at 100 and score clear next steps at 40.

For a new human-authored evaluation, submit the criterion scores against the call and agent. A representative request to the voice QA endpoint looks like this:

```bash theme={null}
curl -X POST https://api.orbit.devotel.io/api/v1/voice/qa-evaluations \
  -H "X-API-Key: $ORBIT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "form_id": "FORM_ID",
    "agent_id": "agent_42",
    "call_id": "call_9f2c",
    "scores": {
      "identity-verified": 100,
      "clear-next-steps": 40
    }
  }'
```

If you are correcting the already-created AI row rather than creating a second evaluation, use the row's **Override** action in the reviewer screen. The Quality API equivalent is `POST /api/v1/quality/evaluations/{evaluation_id}/override`; do not use the create request to duplicate the call's auto-score. The server derives the weighted total from the rubric and enforces the `auto_fail` rule. A response with the evaluation ID and final score is the confirmation to keep for the next step.

The evaluation is linked to the call and its agent. It does not edit the contact's name, phone number, or other contact fields. Where your workspace shows the related call in the contact activity history, the evaluation remains associated through that call rather than becoming a new contact update.

## 4. Follow the agent's acknowledgement or appeal

After the review is saved, the agent sees a notification and the evaluation in their own scorecard/evaluation view. They can open the rubric and call context, then acknowledge the result or appeal it with a written note.

In the dashboard, the agent opens the evaluation, chooses **Appeal**, and explains the disagreement. For this example, the agent writes: “I gave the caller the billing date and said I would send the confirmation after the call; please check the final minute of the recording.” The reviewer returns to the appealed item, checks that segment, and responds by resolving the appeal. Resolution closes the appeal; it does not silently change the grade. If the review changes the score, record that as the appropriate reviewer correction as well.

The API lifecycle uses the evaluation created in the previous step:

```bash theme={null}
# Acknowledge instead of appealing
curl -X POST "https://api.orbit.devotel.io/api/v1/quality/evaluations/$EVALUATION_ID/acknowledge" \
  -H "X-API-Key: $ORBIT_API_KEY"

# Or appeal with a required written note
curl -X POST "https://api.orbit.devotel.io/api/v1/quality/evaluations/$EVALUATION_ID/appeal" \
  -H "X-API-Key: $ORBIT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"appeal_note":"I gave the caller the billing date and said I would send the confirmation after the call; please check the final minute of the recording."}'

# The reviewer resolves the appeal after listening to the cited segment
curl -X POST "https://api.orbit.devotel.io/api/v1/quality/evaluations/$EVALUATION_ID/resolve" \
  -H "X-API-Key: $ORBIT_API_KEY"
```

Only the evaluated agent can acknowledge or appeal; the appeal needs a note, and a reviewer resolves it. Keep the agent's note and reviewer response with the evaluation history so both sides can see how the question was handled.

## 5. Confirm the automatic coaching assignment

The human-reviewed score in this example is **60**, below the configured QA coaching trigger threshold of **70**. After the reviewer finalizes it, open **Voice → Coaching → Coaching plans** and look for the agent's open plan. Its goal includes the source tag `[auto-coaching:qa_score]`, which identifies a QA-score-triggered assignment. The assignment is created from the human-finalized low score; an AI flag by itself is not the coaching trigger.

A low score can result in one open QA coaching plan for that agent rather than a new plan for every later low score while the plan remains open. Attach the coaching action to this existing plan instead of creating a duplicate. The isolated coaching workflow is covered in [Supervisor coaching](/guides/supervisor-coaching).

## 6. Attach practice and close the loop

From the coaching plan, attach a Practice Studio roleplay scenario that rehearses the missed behavior: explain the next step and who owns it. Ask the agent to open **Practice Studio**, complete the scenario as the agent, and review the scenario feedback with the supervisor. Practice is rehearsal, not another call evaluation; it does not erase the original score or add leaderboard points by itself.

On the agent's next graded live call, use the same rubric to check whether the behavior changed. Then review the agent's [scorecard](/guides/my-scorecard) and the [quality leaderboard](/guides/quality-leaderboard). The scorecard reflects the new evaluation, and the leaderboard recomputes from scored calls and its other configured inputs. Compare like-for-like windows and sample sizes; one improved call is evidence to inspect, not proof of a sustained trend.

### Worked example at a glance

| Stage | Example result |
| - | - |
| Rubric | `identity-verified` (required, auto-fail) + `clear-next-steps` (weighted 2×) |
| Recorded call | `call_9f2c`, completed with transcript and assigned agent |
| Grade | Identity 100; next steps 40; derived total about 60/100 |
| Agent response | Appeals with a note pointing to the call's final minute; reviewer checks and resolves |
| Coaching | Open coaching plan for the agent, tagged `[auto-coaching:qa_score]` |

## What good looks like at day 30

Treat these as operating targets for your team, not platform guarantees. Set the baseline from your own eligible call volume and adjust the sample rate if the review workload is too high.

| Measure | Day-30 target | Check |
| - | - | - |
| Eligible calls with an auto-score | At least 90% in the calls you expect the feature to cover | Quality auto-QA coverage and the evaluations list |
| Sampled calls with a completed human review | At least 90% of the sample due during the period | Quality → Evaluations queue |
| Open flagged evaluations older than 2 business days | 0 | Filter flagged items and sort by age |
| Appeals with a reviewer response | 100% within your team's agreed service window | Evaluation appeal history |
| Low-score coaching plans with a next action | 100% have an owner and a practice or coaching step | Voice → Coaching → Coaching plans |
| Repeat-call improvement | Compare the same rubric section over at least 5 later graded calls per coached agent | Agent scorecard and quality leaderboard |

If coverage is low, verify that the auto-score toggle is on and the calls have transcripts, an assigned agent, and an active form. If the queue is growing, reduce the sampling rate or assign review ownership before increasing coverage.

## See also

* [Build a contact-center QA program](/guides/quality-management-program) — rubric authoring, calibration, and evaluation lifecycle depth
* [Supervisor coaching](/guides/supervisor-coaching) — coaching conversations in isolation
* [AI Auto-QA configuration](/guides/qa-autoscore-settings) and [QA sampling](/guides/qa-sampling-settings) — setup and policy detail
* [Agent scorecard](/guides/my-scorecard) and [Quality leaderboard](/guides/quality-leaderboard) — follow scores after the next graded call
* [Practice Studio roleplay](/guides/practice-studio-roleplay) — roleplay setup and scenario walkthrough
* [Quality Management API](/api-reference/quality) — evaluation, claim, override, acknowledgement, and appeal endpoints


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.