Run your first QA evaluation, from rubric to coaching plan
Use one recorded support call to connect the parts of a first-week QA setup: define a small rubric, turn on automated scoring and sampling, review a flagged result, let the agent respond, and follow a low score into coaching. The same rubric stays attached to the call from first score through the coaching follow-up. You need an active voice team, a recorded call with a transcript and an assigned agent, and an owner, admin, or supervisor account for setup and review. The example values below are illustrative; replace IDs with values from your workspace. Auto-score is tenant-controlled and off until you enable it.1. Author one two-section rubric
In the dashboard, open Quality → Rubrics and create Support call – first review. Add two sections:- Required checks: one identity-verification criterion, set as the section’s
auto_failcompliance gate. - Conversation quality: one criterion for clear next steps.
auto_fail flag belongs to the compliance section; the criterion remains individually scoreable.
FORM_ID. Keep this as your active form for the walkthrough: AI Auto-QA uses the newest active evaluation form. The 10% form sample rate is a starting point for a small review sample, not a promise that exactly one of ten calls will appear in every short window.
For the full weighting, calibration, and form lifecycle guidance, see Build a contact-center QA program.
2. Enable auto-score and sampling
Open Settings → QA auto-score, enable the tenant setting, and set the flag threshold to 70. Then open Settings → QA sampling and start with a modest 10% sample. Auto-score evaluates eligible completed calls; sampling provides a separate human-review sample. They work together, but a sampled unscored row and an AI-scored flagged row are different queue items. The first call must be completed, have a transcript, and have an assigned agent. After the next autoscore pass, open Quality → Evaluations and filter for auto-scored or flagged evaluations. Confirm that one row points to the recorded call and showsauto_scored: true. If no row appears, check the auto-score toggle, active form, transcript, agent assignment, and QA auto-score settings. For how the sample share is configured, see QA sampling settings.
For this worked example, the scorer rates identity verified = 100 and clear next steps = 40, for a weighted total of about 60/100. Because 60 is below the 70 flag threshold, the evaluation is marked for a human review. The auto-score is a first pass; the reviewer owns the final human assessment.
3. Review and grade the flagged call
In Quality → Evaluations, select the flagged row for the example call. The reviewer view brings together the call and transcript, the rubric sections and criteria, the automated scores, and the review flag. Listen to the relevant part of the recording before changing a score. In this example, the next step was unclear, so keep identity verification at 100 and score clear next steps at 40. For a new human-authored evaluation, submit the criterion scores against the call and agent. A representative request to the voice QA endpoint looks like this:POST /api/v1/quality/evaluations/{evaluation_id}/override; do not use the create request to duplicate the call’s auto-score. The server derives the weighted total from the rubric and enforces the auto_fail rule. A response with the evaluation ID and final score is the confirmation to keep for the next step.
The evaluation is linked to the call and its agent. It does not edit the contact’s name, phone number, or other contact fields. Where your workspace shows the related call in the contact activity history, the evaluation remains associated through that call rather than becoming a new contact update.
4. Follow the agent’s acknowledgement or appeal
After the review is saved, the agent sees a notification and the evaluation in their own scorecard/evaluation view. They can open the rubric and call context, then acknowledge the result or appeal it with a written note. In the dashboard, the agent opens the evaluation, chooses Appeal, and explains the disagreement. For this example, the agent writes: “I gave the caller the billing date and said I would send the confirmation after the call; please check the final minute of the recording.” The reviewer returns to the appealed item, checks that segment, and responds by resolving the appeal. Resolution closes the appeal; it does not silently change the grade. If the review changes the score, record that as the appropriate reviewer correction as well. The API lifecycle uses the evaluation created in the previous step:5. Confirm the automatic coaching assignment
The human-reviewed score in this example is 60, below the configured QA coaching trigger threshold of 70. After the reviewer finalizes it, open Voice → Coaching → Coaching plans and look for the agent’s open plan. Its goal includes the source tag[auto-coaching:qa_score], which identifies a QA-score-triggered assignment. The assignment is created from the human-finalized low score; an AI flag by itself is not the coaching trigger.
A low score can result in one open QA coaching plan for that agent rather than a new plan for every later low score while the plan remains open. Attach the coaching action to this existing plan instead of creating a duplicate. The isolated coaching workflow is covered in Supervisor coaching.
6. Attach practice and close the loop
From the coaching plan, attach a Practice Studio roleplay scenario that rehearses the missed behavior: explain the next step and who owns it. Ask the agent to open Practice Studio, complete the scenario as the agent, and review the scenario feedback with the supervisor. Practice is rehearsal, not another call evaluation; it does not erase the original score or add leaderboard points by itself. On the agent’s next graded live call, use the same rubric to check whether the behavior changed. Then review the agent’s scorecard and the quality leaderboard. The scorecard reflects the new evaluation, and the leaderboard recomputes from scored calls and its other configured inputs. Compare like-for-like windows and sample sizes; one improved call is evidence to inspect, not proof of a sustained trend.Worked example at a glance
What good looks like at day 30
Treat these as operating targets for your team, not platform guarantees. Set the baseline from your own eligible call volume and adjust the sample rate if the review workload is too high.
If coverage is low, verify that the auto-score toggle is on and the calls have transcripts, an assigned agent, and an active form. If the queue is growing, reduce the sampling rate or assign review ownership before increasing coverage.
See also
- Build a contact-center QA program — rubric authoring, calibration, and evaluation lifecycle depth
- Supervisor coaching — coaching conversations in isolation
- AI Auto-QA configuration and QA sampling — setup and policy detail
- Agent scorecard and Quality leaderboard — follow scores after the next graded call
- Practice Studio roleplay — roleplay setup and scenario walkthrough
- Quality Management API — evaluation, claim, override, acknowledgement, and appeal endpoints