Settings → QA Autoscore console
Open Settings → QA Autoscore at/settings/qa-autoscore. This console is the control room for automated LLM-judged QA scoring: it holds the rule threshold editor, the calibration workflow that tests a rule against real conversation samples, and the readiness panel that tells you whether your rules are safe to activate.
The console is gated to owner and admin roles. Supervisors can read the readiness panel on the Quality hub but cannot create or calibrate autoscore rules.
Every control on this page is tenant-owned: your organization chooses which rubrics to auto-score, what threshold to enforce, and when a rule is ready for production. The AI model scores against your own active evaluation form — it never invents a rubric.
1. What QA Autoscore is and who needs it
QA Autoscore is the tenant’s launch gate for LLM-judged evaluation. Instead of a human reviewer grading every call transcript against your scorecard form, you write one or more autoscore rules — each rule targets a section or criterion from your evaluation form, defines a pass/fail threshold, and, once calibrated and confirmed ready, automatically scores every eligible conversation the sweep processes. You need this console if:- Your QA program processes more calls than human evaluators can manually review each week.
- You want an AI pre-screen that flags borderline conversations before a human sees them, so evaluator time is spent on the ambiguous calls rather than the obviously passing ones.
- You want per-rubric scoring that gives you coverage across every evaluation criterion, not just an aggregate total.
- Your QA program is small enough that human review covers every conversation.
- You only use the global AI Auto-QA toggle (the master on/off and single flag threshold under the AI Auto-QA configuration guide). This console adds per-rubric rules and calibration on top of that global switch.
Where this console sits in the QA pipeline
Three surfaces touch QA scoring, and they are easy to confuse:
The global toggle gates the entire sweep; the per-rubric rules on this console decide what the sweep scores and how it judges each rubric.
2. The readiness panel
The readiness panel is the top card on the console. It pre-flights the entire autoscore pipeline before a rule goes live and answers one question: “If I activate this rule, will it score correctly?”What the panel measures
When all four blocks pass, the panel shows a green Ready for go-live badge. When any block fails, the panel renders one chip per failed gate — each chip links to the surface that fixes it.
Reading a readiness score below 70
A readiness score below 70 means the rule’s calibration run showed one of two problems:- Low agreement with human scores. The AI scorecard and the human-assigned scorecard disagree too often, and the rule’s threshold is either too strict (failing conversations a human would pass) or too loose (passing conversations a human would fail).
- Insufficient calibration data. The sample used for calibration is too small. Re-run calibration against a larger sample — at least 20 conversations — and re-check.
3. The calibration workflow
Calibration is the process of running an autoscore rule against a sample of already-human-scored conversations and comparing the AI output to the human score — so you know, before go-live, whether the rule agrees with your evaluators.Step 1 — select a rule to calibrate
In the Rules tab of the QA Autoscore console, pick the rule you want to test. If no rule exists yet, create one first:- Click Create rule.
- Pick the rubric criterion from your active evaluation form — each criterion maps to one section or weighted item on the form.
- Set the pass threshold (0–100). Conversations scored at or above this value by the model are marked passing; those below are flagged for review.
- Save the rule. It is now listed in the rules table, ready for calibration.
Step 2 — pick a calibration sample
Click Calibrate on the rule row. The calibration panel opens and prompts you to pick a sample source:- Human-scored conversations (recommended). The calibration run compares the AI score against the existing human score on each conversation. This is the ground-truth path — you need at least 10 human-scored calls for the agreement metric to be meaningful, and 20 or more for a reliable readiness score.
- Unscored conversations from the last N days. The run scores the conversations with the model and reports the distribution, but without a ground-truth comparison the agreement metric is unavailable and the readiness score will be capped.
Step 3 — run calibration
Click Run calibration. The console sends the selected conversations to the model with your evaluation form’s rubric for the chosen criterion and the rule’s threshold. Each conversation gets an AI scorecard for that criterion. The run takes up to 60 seconds depending on sample size. The panel shows progress per conversation; do not close the tab while the run is in flight.Step 4 — read the calibration report
When the run finishes, the panel renders the calibration report:
Use the report to decide whether the rule is ready. Adjust the threshold or reword the rubric criterion in the evaluation form, then re-run calibration until the agreement rate and the false-positive/false-negative balance are acceptable.
4. Rule thresholds and per-rubric overrides
Each autoscore rule carries one threshold: a score from 0 to 100. A conversation scored at or above the threshold is marked passing; a conversation scored below is flagged.The global flag threshold vs. per-rule thresholds
The global flag threshold (set on the AI Auto-QA tab) is a fallback — it applies to any rubric criterion that does not have its own autoscore rule. A per-rule threshold on this console overrides the global threshold for that specific criterion.Tuning a threshold from the calibration report
The calibration report gives you the data to tune:- False positives are high. Raise the threshold — the AI is letting borderline conversations through.
- False negatives are high. Lower the threshold — the AI is flagging conversations a human would pass, and your review queue will balloon.
- Agreement rate is low with no clear bias. The rubric wording may be ambiguous. Open the evaluation form in Quality → Evaluation forms and tighten the criterion description. A rubric that reads “Agent was polite” is harder to score consistently than “Agent used the customer’s name, acknowledged the issue, and offered a clear next step.”
5. How autoscore interacts with QA Sampling
QA Autoscore and QA Sampling are complementary but independent:
The sampling gate decides what enters scoring for human review, but it does not affect autoscore. Autoscore processes every eligible call regardless of whether it was sampled for human review. The two pipelines write into the same
qa_evaluations table, distinguished by auto_scored: true (AI) vs. the human reviewer’s ID.
A common setup: QA Sampling assigns 3 conversations per agent per week to human evaluators, while QA Autoscore scores every call against a per-rubric rule set, flagging only the borderline ones. The evaluator then reconciles the flagged AI scores against their own manual scores, using the AI pre-screen to focus on the calls that need attention.
6. Worked example: create an autoscore rule, calibrate, and confirm readiness
A contact center wants to auto-score every call against their “Issue Resolution” rubric criterion, with a threshold of 75 — conversations the AI scores below 75 on resolution should be flagged for human review.Step 1 — verify prerequisites
Open Settings → AI Auto-QA and confirm the global toggle is on. Open Quality → Evaluation forms and confirm at least one form is active. Both gates must be green on the readiness panel.Step 2 — create the rule
- In Settings → QA Autoscore, open the Rules tab.
- Click Create rule.
- Select Issue Resolution from the rubric criterion picker. The picker lists every criterion from your active evaluation form.
- Set the pass threshold to 75.
- Save. The rule appears in the rules table.
Step 3 — calibrate
- Click Calibrate on the Issue Resolution rule row.
- Select Human-scored conversations as the sample source.
- Set the sample size to 30 (the minimum for a reliable agreement metric).
- Click Run calibration.
- Wait for the run to complete. The panel shows progress — do not close the tab.
Step 4 — read the report
The calibration report shows:
The agreement rate (88%) is above the 85% target. The false-negative rate (10%) is acceptable — only 3 out of 30 conversations would be flagged unnecessarily. The readiness panel shows a readiness score of 88 and a green Ready for go-live badge.
Step 5 — confirm and go live
The rule is calibrated and ready. The next sweep tick picks it up, and every eligible conversation starts receiving an AI scorecard for Issue Resolution. Flagged conversations appear in the Quality → Evaluations “Needs human review” queue.Step 6 — monitor the first week
After one week, return to the readiness panel and check:- The coverage gauge on the Quality hub — is the share of eligible calls being scored holding above 90%?
- The flagged share — are you getting the expected volume of flagged rows?
- Spot-check five flagged rows against your own evaluation — does the AI verdict hold up?
7. Edge cases
Calibration drift
Over time, the AI model and your evaluation form can drift apart — the form’s rubric wording changes, the model is updated, or your evaluators’ standards evolve. Calibration is not a one-time event. Re-run calibration when:- You edit the evaluation form’s rubric criteria.
- The agreement rate on spot-checked flagged rows drops below your target.
- You add a new agent cohort or team whose call patterns differ.
- A new model version is deployed (the model version used for scoring is noted in the sweep’s metadata).
Scope: inbox vs. organization
Autoscore rules are organization-scoped. Every eligible call in the organization is scored against every active rule. There is no per-queue, per-agent, or per-inbox rule scope. If you need different thresholds for different teams, create separate rules targeting the same rubric criterion but assign them to different agent groups through the rule scope picker (available on the rule editor when agent-group segmentation is enabled). A rule scoped to a group only scores calls handled by agents in that group.Per-rubric overrides and the global threshold
If you define a rule for “Tone and Empathy” with threshold 85, and the global flag threshold is 70, then:- Tone and Empathy uses 85 — the per-rule override.
- Every other criterion without a rule uses 70 — the global fallback.
- If you later delete the Tone and Empathy rule, that criterion falls back to 70 on the next sweep tick.
Multiple rules for the same criterion
The console blocks creating a second rule for a rubric criterion that already has one. If you need a different threshold for the same criterion, edit the existing rule, or scope it to a different agent group.8. Troubleshooting
The readiness panel shows “No active evaluation form.” This gate is not specific to autoscore rules — it is the same gate the global AI Auto-QA sweep checks. Open Quality → Evaluation forms, create or re-enable a form, and the gate clears on the next refresh. Calibration run shows 0% coverage. The sample conversations do not have transcripts, or the rubric criterion does not apply to them. Confirm transcripts are being extracted for completed calls, and check that the criterion exists in your active evaluation form. Agreement rate is stuck below 70% after tuning the threshold. The rubric wording may be the blocker, not the threshold. Open the evaluation form, read the criterion description as the model would read it, and ask: “Could two people read this and give different scores?” If the answer is yes, rewrite the criterion to be specific about the observable behaviour: replace “Agent resolved the issue” with “Agent confirmed the resolution with the customer, stated the action taken, and asked if anything else was needed.” A rule is active but conversations are not being scored. Walk the gates in order: global toggle on → active form exists → rule is saved and active → conversation has a transcript and an assigned agent. Pure-IVR or bot-only calls are skipped — they have no agent to score. The calibration sample picker is empty. Your organization has no human-scored conversations in the lookback window. Complete a few manual evaluations first, or switch to unscored conversations as the sample source — the agreement metric will be unavailable, but you can still read the score distribution.See also
- AI Auto-QA configuration — the global toggle, flag threshold, and sweep lifecycle.
- AI auto-QA readiness card — the dashboard card that pre-flights the pipeline.
- QA Sampling settings — the sampling gate that decides how many conversations get human-reviewed.
- Quality Management Program overview — where this console fits into your QA program.
- Quality Management API reference — the full endpoint and filter matrix.
- Settings hub — every settings console in one map.