Call QA scorecards and auto-scoring
Quality monitoring is the part of a contact-centre stack that usually means buying a separate per-seat module. Orbit gives you the full QA loop where the calls already are — grade any recorded call against a deterministic rubric, follow each agent’s rolling composite, find the recordings that matter by keyword instead of scrubbing audio, and let the AI auto-score every call so a supervisor only spot-audits a sample. This page covers the supervisor-authored scorecard surface (the inline rubric on the call detail page and its API), the per-agent trend view, the transcript and voicemail search library, and the auto-scoring scheduler that grades calls against your tenant’s QA evaluation form. For the richer conversation-grading surface with weighted custom scorecards, appeals, and calibration sessions, see the Quality → QA Evaluations dashboard instead — this page is the call-level rubric that lives on the call detail page. Base path:/api/v1/voice
Authentication: Clerk session (Authorization: Bearer <token>) or API key (X-API-Key).
Scope: voice read for reads; the scorecard upsert is restricted to owners and admins (the dashboard’s supervisor persona).
The rubric: five dimensions, 1–5 each
A scorecard grades one call on five dimensions, each on a 1–5 integer scale with an optional per-dimension note (up to 2,000 characters) and an optional overall note (up to 4,000 characters):
The reviewer is always the authenticated supervisor making the call — the API never trusts a client-supplied reviewer identity, and every save is written to the audit log with the five dimensions so a quarterly QA review can replay who graded what and when.
Composite is the equal-weighted mean of the five scores — the headline number a team tracks week-over-week.
Where it lives
- Call detail page. Open any call under Voice → Calls → your-call. Owners and admins see an inline editor (rubric chips + notes fields). Everyone else with call access sees the persisted rubric read-only — the editor never flashes for non-admins because the backend enforces the same gate.
- Agent trend. Below the editor, the same panel charts the answering agent’s daily composite average over a 30-day window (customisable 1–90 days), so a supervisor can see a single save move the agent’s rolling quality number. Agents can’t be mis-attributed — the backend resolves the answering agent from the call’s metadata (
agent_user_id, fallbackuser_id), not from a client-provided field. - Recording library. The unified QA worklist — every recording, its diarised transcript, the latest linked evaluation score, and the recording QC verdict in one query — is documented on Searchable recording and transcript library (QA).
Endpoints
The save body accepts each
<dimension>_score (integer 1–5, required) plus <dimension>_notes / overall_notes (optional). Re-saving the same call updates the rubric in place and bumps updated_at — the upsert key is the call id, so two supervisors re-reviewing the same call converge on one canonical scorecard rather than forking. The first save returns 201; a re-save returns 200.
The trend endpoint rolls the scorecards up server-side into daily buckets:
avg_composite is the equal-weighted mean of the five avg_* per-dimension averages — no client-side math needed. Reading a call that has not been reviewed returns 404 with a clear message rather than a generic error, so the dashboard renders “Not yet reviewed” cleanly on a fresh tenant.
Worked examples
The sections above name the operations; these four runnable loops cover the exact request and response shapes so you can drive the QA loop from a script: upsert a five-dimension scorecard, read it back, derive an agent’s composite trend, and handle the “not yet reviewed” branch cleanly. Every sample below uses the placeholder tokenscall_8fb2d1b7 for the call id, qasc_7d3e2a1b9c4f for the scorecard id, and user_0q2k9 / user_9z4p1 for the reviewer and agent — substitute your real ids. Sandbox keys only: dv_test_sk_YOUR_KEY for curl, process.env.ORBIT_API_KEY in code.
1. Upsert the five-dimension scorecard
The reviewer id is never client-supplied — the API stamps the authenticated supervisor from the session, so this loop keeps to the rubric scores and notes only.201; re-saving the same call updates the rubric in place (the upsert key is the call id) and returns 200. The saved row echoes back:
agent_user_id is resolved server-side from the call’s metadata (the answering agent for inbound, the dialing agent for outbound), never from the request body.
2. Read one call’s scorecard
3. Derive the agent’s composite trend
The trend loop stands alone as curl — one endpoint, one envelope.avg_composite for each bucket is server-side math: the server averages each call’s per-call composite — the equal-weighted mean of that call’s five dimension scores — then rounds to two decimals. In the 2026-09-23 bucket above, four calls averaged (4.25 + 3.75 + 4.0 + 4.25 + 4.0) ÷ 5 = 4.05 per call. A window with no scorecards returns rows: [], never an error, so a “no reviews yet” branch renders cleanly on a fresh tenant.
4. Handle “Not yet reviewed”
Reading an unreviewed call returns404 with a clear envelope so your UI shows a “Not yet reviewed” state instead of a generic error:
404 / NOT_FOUND pair — the dashboard’s “Nothing reviewed yet” queue and the SDK’s OrbitError both key off it.
Finding the right call: searchable recordings and transcripts
You no longer scrub audio to find the call a QA review needs. Three surfaces index every spoken word into full-text search so you can find recordings by keyword, phrase, or negation:- Call transcript search —
GET /api/v1/voice/calls/searchsearches the full post-call speech-to-text transcript stored oncall_logs.metadata.transcript/transcript_data.transcript. Query params:q(required, up to 500 chars),limit(default 25, max 100),since/until(ISO time range),direction(inbound/outbound). - Voicemail transcript search —
GET /api/v1/voice/voicemails/searchtargets voicemail messages specifically — same query shape, narrower surface (q,limitup to 50,since,until). - PII redaction applies. Tenants with transcript redaction enabled can still search — the snippet in each result respects the tenant’s redaction mode, so raw caller speech never leaks through a QA surface.
billing OR refund, "angry customer" (exact phrase), and cancellation -reschedule (with an exclusion) all work. Each result carries the call or message id, the standard from/to/direction columns, a relevance rank, and a ts_headline snippet with <mark> tags around matched tokens, ready for inline highlight in a dashboard list.
For the unified QA worklist — every recording class, the finalised diarised transcript segments, the latest linked evaluation score, and the recording QC verdict in one keyset-paginated surface — use the recording library: Searchable recording and transcript library (QA).
Auto-scoring: AI grades every call against your QA form
Manual rubric grading covers the calls a supervisor picks. Auto-scoring covers the other 100% — a scheduler grades every completed call against your tenant’s active QA evaluation form (the same weighted scorecards you build in Quality → QA Evaluations) as soon as the transcript lands. Supervisors then spend their time spot-auditing the AI’s work rather than scoring every call from scratch. How it works:- A background sweep (60-second cadence) picks up completed calls that have a transcript and no auto-scored evaluation yet.
- The transcript (capped at 12,000 characters ≈ 3,000 tokens per call) is scored against each criterion on your active QA evaluation form. The LLM only scores individual criteria on each criterion’s own 0..max_score scale; the weighted rollup is computed server-side deterministically, so a form’s section weights and auto-fail sections can’t be undermined by LLM arithmetic.
- The result is written as a normal QA evaluation with the
auto-llmreviewer id andauto_scored = trueprovenance, so downstream readers and your own “percent auto-scored” reports can tell it from a human-authored score. - Scores below a flag-for-review threshold (default 70 out of 100, comparable to the vendor defaults) are marked
flagged_for_reviewso the supervisor’s “needs human review” queue carries only the borderline calls — the Five9 School of Thought QA pattern.
qa_autoscore (set enabled: true, optionally flag_threshold: 0..100), or through the dashboard in Settings → QA Auto-Score. A per-tenant cap and per-call LLM deadline keep grading latency under ~2 minutes from hangup without ever burning the model budget on a stuck call.
Auto-scored evaluations read exactly like human-authored ones to the QA Evaluations dashboard — the agent can still acknowledge or appeal them, and a supervisor can override the score with a normal human review.
Related
- Real-time agent-assist whisper coaching — the live mid-call counterpart to this post-call grading loop: suggested replies, KB articles, and a rolling sentiment read streamed to the agent and supervisor wallboard
- Searchable recording and transcript library (QA) — the unified QA worklist for recordings, transcripts, and linked evaluation scores