> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Settings → QA Autoscore console: readiness panel, calibration, and rule thresholds

> Operate the /settings/qa-autoscore console end to end — read the readiness panel, calibrate an autoscore rule against a sample of conversations, set per-rubric thresholds, and confirm readiness before go-live.

# Settings → QA Autoscore console

Open **Settings → QA Autoscore** at `/settings/qa-autoscore`. This console is the control room for automated LLM-judged QA scoring: it holds the rule threshold editor, the calibration workflow that tests a rule against real conversation samples, and the readiness panel that tells you whether your rules are safe to activate.

The console is gated to **owner** and **admin** roles. Supervisors can read the readiness panel on the Quality hub but cannot create or calibrate autoscore rules.

<Note>
  Every control on this page is **tenant-owned**: your organization chooses which rubrics to auto-score, what threshold to enforce, and when a rule is ready for production. The AI model scores against your own active evaluation form — it never invents a rubric.
</Note>

***

## 1. What QA Autoscore is and who needs it

QA Autoscore is the tenant's launch gate for **LLM-judged evaluation**. Instead of a human reviewer grading every call transcript against your scorecard form, you write one or more **autoscore rules** — each rule targets a section or criterion from your evaluation form, defines a pass/fail threshold, and, once calibrated and confirmed ready, automatically scores every eligible conversation the sweep processes.

**You need this console if:**

* Your QA program processes more calls than human evaluators can manually review each week.
* You want an AI pre-screen that flags borderline conversations before a human sees them, so evaluator time is spent on the ambiguous calls rather than the obviously passing ones.
* You want per-rubric scoring that gives you coverage across every evaluation criterion, not just an aggregate total.

**You do not need this console if:**

* Your QA program is small enough that human review covers every conversation.
* You only use the global AI Auto-QA toggle (the master on/off and single flag threshold under the [AI Auto-QA configuration](/guides/qa-autoscore-settings) guide). This console adds per-rubric rules and calibration on top of that global switch.

### Where this console sits in the QA pipeline

Three surfaces touch QA scoring, and they are easy to confuse:

| Surface | Route | What it does |
| - | - | - |
| **AI Auto-QA toggle** | `/settings/qa-autoscore` (global tab) | Master enable/disable, global flag threshold, sweep configuration. One on/off for the whole tenant. |
| **QA Autoscore console** | `/settings/qa-autoscore` (rules tab) | Per-rubric rules, calibration workflow, readiness panel. This is the surface this guide covers. |
| **QA Sampling** | `/settings/qa-sampling` | Decides which conversations enter the human-review queue — the sampling gate sits before scoring, so a conversation must be sampled before a human sees it. Autoscore operates independently of sampling. |

The global toggle gates the entire sweep; the per-rubric rules on this console decide *what* the sweep scores and *how* it judges each rubric.

***

## 2. The readiness panel

The readiness panel is the top card on the console. It pre-flights the entire autoscore pipeline before a rule goes live and answers one question: **"If I activate this rule, will it score correctly?"**

### What the panel measures

| Block | What it reports | Green threshold |
| - | - | - |
| **Generation gate** | Global AI Auto-QA toggle is on, and an active evaluation form exists. | Both pass. |
| **Rule gate** | At least one autoscore rule is defined and saved. | At least one rule exists. |
| **Calibration status** | The selected rule has been calibrated against a sample of conversations, and the calibration run produced a score distribution. | Calibration run exists and is recent. |
| **Readiness score** | A composite 0–100 score computed from calibration results: agreement rate between the rule and the ground-truth human score on the sample, plus coverage of the rubric criteria. | ≥ 70 |

When all four blocks pass, the panel shows a green **Ready for go-live** badge. When any block fails, the panel renders one chip per failed gate — each chip links to the surface that fixes it.

### Reading a readiness score below 70

A readiness score below 70 means the rule's calibration run showed one of two problems:

* **Low agreement with human scores.** The AI scorecard and the human-assigned scorecard disagree too often, and the rule's threshold is either too strict (failing conversations a human would pass) or too loose (passing conversations a human would fail).
* **Insufficient calibration data.** The sample used for calibration is too small. Re-run calibration against a larger sample — at least 20 conversations — and re-check.

The readiness score is advisory, not a hard gate. You can activate a rule with a score below 70, but a low readiness score means the AI judgement is likely to diverge from your human evaluators, and flagged rows may need more frequent overrides.

***

## 3. The calibration workflow

Calibration is the process of running an autoscore rule against a sample of already-human-scored conversations and comparing the AI output to the human score — so you know, before go-live, whether the rule agrees with your evaluators.

### Step 1 — select a rule to calibrate

In the **Rules** tab of the QA Autoscore console, pick the rule you want to test. If no rule exists yet, create one first:

1. Click **Create rule**.
2. Pick the **rubric criterion** from your active evaluation form — each criterion maps to one section or weighted item on the form.
3. Set the **pass threshold** (0–100). Conversations scored at or above this value by the model are marked passing; those below are flagged for review.
4. Save the rule. It is now listed in the rules table, ready for calibration.

You can define as many rules as you have rubric criteria. Each rule scores independently, and the readiness panel runs per rule.

### Step 2 — pick a calibration sample

Click **Calibrate** on the rule row. The calibration panel opens and prompts you to pick a sample source:

* **Human-scored conversations** (recommended). The calibration run compares the AI score against the existing human score on each conversation. This is the ground-truth path — you need at least 10 human-scored calls for the agreement metric to be meaningful, and 20 or more for a reliable readiness score.
* **Unscored conversations from the last N days.** The run scores the conversations with the model and reports the distribution, but without a ground-truth comparison the agreement metric is unavailable and the readiness score will be capped.

Select the sample size. The minimum is 10; 20–50 gives a reliable distribution. The panel pre-selects a random draw from the eligible pool — you can narrow by agent, date range, or conversation tag before confirming.

### Step 3 — run calibration

Click **Run calibration**. The console sends the selected conversations to the model with your evaluation form's rubric for the chosen criterion and the rule's threshold. Each conversation gets an AI scorecard for that criterion.

The run takes up to 60 seconds depending on sample size. The panel shows progress per conversation; do not close the tab while the run is in flight.

### Step 4 — read the calibration report

When the run finishes, the panel renders the calibration report:

| Metric | What it means |
| - | - |
| **Agreement rate** | Share of conversations where the AI pass/fail verdict matches the human verdict. A rate of 85%+ is strong; below 70% suggests the threshold or the rubric wording needs tuning. |
| **Mean AI score vs. mean human score** | The average scores side by side. A gap larger than 10 points suggests systematic bias in one direction — the AI is either consistently stricter or consistently more lenient than your evaluators. |
| **False positive rate** | Conversations the AI passed but a human failed. A high false-positive rate means the threshold is too loose. |
| **False negative rate** | Conversations the AI failed but a human passed. A high false-negative rate means the threshold is too strict and your review queue will fill with conversations that do not need attention. |
| **Coverage** | Share of the calibration sample that the rule was able to score. Some conversations may be unscorable — missing transcripts, calls with no agent, or rubric criteria that do not apply. |

Use the report to decide whether the rule is ready. Adjust the threshold or reword the rubric criterion in the evaluation form, then re-run calibration until the agreement rate and the false-positive/false-negative balance are acceptable.

***

## 4. Rule thresholds and per-rubric overrides

Each autoscore rule carries one threshold: a score from 0 to 100. A conversation scored at or above the threshold is marked passing; a conversation scored below is flagged.

### The global flag threshold vs. per-rule thresholds

The global flag threshold (set on the **AI Auto-QA** tab) is a fallback — it applies to any rubric criterion that does not have its own autoscore rule. A per-rule threshold on this console **overrides** the global threshold for that specific criterion.

| Scenario | Which threshold applies |
| - | - |
| You have a rule for "Tone and Empathy" with threshold 80. | 80 applies for that criterion. |
| You have no rule for "Resolution Completeness." | The global flag threshold applies. |
| You disable a per-rule threshold by setting it to "inherit." | The global flag threshold applies. |

### Tuning a threshold from the calibration report

The calibration report gives you the data to tune:

* **False positives are high.** Raise the threshold — the AI is letting borderline conversations through.
* **False negatives are high.** Lower the threshold — the AI is flagging conversations a human would pass, and your review queue will balloon.
* **Agreement rate is low with no clear bias.** The rubric wording may be ambiguous. Open the evaluation form in **Quality → Evaluation forms** and tighten the criterion description. A rubric that reads "Agent was polite" is harder to score consistently than "Agent used the customer's name, acknowledged the issue, and offered a clear next step."

Save the rule after adjusting the threshold; the change takes effect on the sweep's next tick.

***

## 5. How autoscore interacts with QA Sampling

QA Autoscore and QA Sampling are complementary but independent:

| | QA Autoscore | QA Sampling |
| - | - | - |
| **What it decides** | Whether a conversation is auto-scored by the model, and whether it is flagged for human review. | Which conversations a human evaluator sees for manual scoring. |
| **Who it serves** | The AI pipeline — it writes auto-scored rows into the QA ledger. | Human evaluators — it assigns review work to their inboxes. |
| **Configuration surface** | `/settings/qa-autoscore` | `/settings/qa-sampling` |
| **Can they overlap?** | Yes — the same conversation can carry both an AI scorecard (from autoscore) and a human scorecard (from sampling), and the supervisor reconciles the two in the evaluation form UI. | |

The **sampling gate decides what enters scoring** for human review, but it does not affect autoscore. Autoscore processes every eligible call regardless of whether it was sampled for human review. The two pipelines write into the same `qa_evaluations` table, distinguished by `auto_scored: true` (AI) vs. the human reviewer's ID.

A common setup: QA Sampling assigns 3 conversations per agent per week to human evaluators, while QA Autoscore scores every call against a per-rubric rule set, flagging only the borderline ones. The evaluator then reconciles the flagged AI scores against their own manual scores, using the AI pre-screen to focus on the calls that need attention.

***

## 6. Worked example: create an autoscore rule, calibrate, and confirm readiness

A contact center wants to auto-score every call against their "Issue Resolution" rubric criterion, with a threshold of 75 — conversations the AI scores below 75 on resolution should be flagged for human review.

### Step 1 — verify prerequisites

Open **Settings → AI Auto-QA** and confirm the global toggle is on. Open **Quality → Evaluation forms** and confirm at least one form is active. Both gates must be green on the readiness panel.

### Step 2 — create the rule

1. In **Settings → QA Autoscore**, open the **Rules** tab.
2. Click **Create rule**.
3. Select **Issue Resolution** from the rubric criterion picker. The picker lists every criterion from your active evaluation form.
4. Set the **pass threshold** to **75**.
5. Save. The rule appears in the rules table.

### Step 3 — calibrate

1. Click **Calibrate** on the Issue Resolution rule row.
2. Select **Human-scored conversations** as the sample source.
3. Set the sample size to **30** (the minimum for a reliable agreement metric).
4. Click **Run calibration**.
5. Wait for the run to complete. The panel shows progress — do not close the tab.

### Step 4 — read the report

The calibration report shows:

| Metric | Value |
| - | - |
| Agreement rate | 88% |
| Mean AI score | 82 |
| Mean human score | 79 |
| False positive rate | 7% |
| False negative rate | 10% |
| Coverage | 100% |

The agreement rate (88%) is above the 85% target. The false-negative rate (10%) is acceptable — only 3 out of 30 conversations would be flagged unnecessarily. The readiness panel shows a readiness score of **88** and a green **Ready for go-live** badge.

### Step 5 — confirm and go live

The rule is calibrated and ready. The next sweep tick picks it up, and every eligible conversation starts receiving an AI scorecard for Issue Resolution. Flagged conversations appear in the **Quality → Evaluations** "Needs human review" queue.

### Step 6 — monitor the first week

After one week, return to the readiness panel and check:

* The **coverage gauge** on the Quality hub — is the share of eligible calls being scored holding above 90%?
* The **flagged share** — are you getting the expected volume of flagged rows?
* Spot-check five flagged rows against your own evaluation — does the AI verdict hold up?

If the flagged share is too high, raise the threshold. If agreement drifts (e.g., evaluators are overturning too many AI decisions), re-run calibration against a fresh sample and tune.

***

## 7. Edge cases

### Calibration drift

Over time, the AI model and your evaluation form can drift apart — the form's rubric wording changes, the model is updated, or your evaluators' standards evolve. Calibration is not a one-time event.

**Re-run calibration when:**

* You edit the evaluation form's rubric criteria.
* The agreement rate on spot-checked flagged rows drops below your target.
* You add a new agent cohort or team whose call patterns differ.
* A new model version is deployed (the model version used for scoring is noted in the sweep's metadata).

### Scope: inbox vs. organization

Autoscore rules are **organization-scoped**. Every eligible call in the organization is scored against every active rule. There is no per-queue, per-agent, or per-inbox rule scope.

If you need different thresholds for different teams, create separate rules targeting the same rubric criterion but assign them to different agent groups through the **rule scope** picker (available on the rule editor when agent-group segmentation is enabled). A rule scoped to a group only scores calls handled by agents in that group.

### Per-rubric overrides and the global threshold

If you define a rule for "Tone and Empathy" with threshold 85, and the global flag threshold is 70, then:

* Tone and Empathy uses 85 — the per-rule override.
* Every other criterion without a rule uses 70 — the global fallback.
* If you later delete the Tone and Empathy rule, that criterion falls back to 70 on the next sweep tick.

### Multiple rules for the same criterion

The console blocks creating a second rule for a rubric criterion that already has one. If you need a different threshold for the same criterion, edit the existing rule, or scope it to a different agent group.

***

## 8. Troubleshooting

**The readiness panel shows "No active evaluation form."** This gate is not specific to autoscore rules — it is the same gate the global AI Auto-QA sweep checks. Open **Quality → Evaluation forms**, create or re-enable a form, and the gate clears on the next refresh.

**Calibration run shows 0% coverage.** The sample conversations do not have transcripts, or the rubric criterion does not apply to them. Confirm transcripts are being extracted for completed calls, and check that the criterion exists in your active evaluation form.

**Agreement rate is stuck below 70% after tuning the threshold.** The rubric wording may be the blocker, not the threshold. Open the evaluation form, read the criterion description as the model would read it, and ask: "Could two people read this and give different scores?" If the answer is yes, rewrite the criterion to be specific about the observable behaviour: replace "Agent resolved the issue" with "Agent confirmed the resolution with the customer, stated the action taken, and asked if anything else was needed."

**A rule is active but conversations are not being scored.** Walk the gates in order: global toggle on → active form exists → rule is saved and active → conversation has a transcript and an assigned agent. Pure-IVR or bot-only calls are skipped — they have no agent to score.

**The calibration sample picker is empty.** Your organization has no human-scored conversations in the lookback window. Complete a few manual evaluations first, or switch to unscored conversations as the sample source — the agreement metric will be unavailable, but you can still read the score distribution.

***

## See also

* [AI Auto-QA configuration](/guides/qa-autoscore-settings) — the global toggle, flag threshold, and sweep lifecycle.
* [AI auto-QA readiness card](/guides/qa-autoscore-readiness-panel) — the dashboard card that pre-flights the pipeline.
* [QA Sampling settings](/guides/qa-sampling-settings) — the sampling gate that decides how many conversations get human-reviewed.
* [Quality Management Program overview](/guides/quality-management-program) — where this console fits into your QA program.
* [Quality Management API reference](/api-reference/quality) — the full endpoint and filter matrix.
* [Settings hub](/settings/overview) — every settings console in one map.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.