> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# The LLM-judged evaluation ledger

> How the Quality → Evaluations page records every rubric judgment — LLM-judged, human review, auto-sampled, or CSAT-triggered — as one per-call ledger row, and how a rubric pass/fail decision moves downstream to coaching and the leaderboard.

# The LLM-judged evaluation ledger

Read the [call-quality evaluation lifecycle](/concepts/quality-evaluation-lifecycle) for the operations layer around scoring; read [QA evaluations and the performance leaderboard](/concepts/qa-leaderboard-and-evaluations) for the evaluation form and leaderboard contract. This page names the thing both pipeline concepts leave implicit: the **evaluation ledger** — the per-call judgment record the **Quality → Evaluations** page renders, where every row carries *who judged it* (an LLM scorer, a human reviewer, or a queued placeholder), *what triggered the judgment*, and *where the pass/fail decision goes next*.

## Why the judge matters, not just the score

A leaderboard compresses each agent to a QA average. When that average drops, the supervisor's first question is never "how low" — it is "who said so, and which to trust." Different kinds of judges produce different trust levels:

* A row authored by a **human reviewer** is an official audit-grade judgment with an acknowledge/appeal lifecycle.
* A row authored by the **auto-scoring LLM judge** is a machine judgment — correctable by a human override, but until overridden it counts toward the official average exactly like a human score.
* A **queued placeholder** — an auto-sampling or CSAT-trigger row with a zero pending score — is not a judgment at all and is excluded from every average until a human claims and scores it.
* A **self-evaluation** — the agent scoring their own call — is coaching material, deliberately excluded from official averages and the leaderboard so agents cannot inflate their own rank.

The average alone cannot answer any of that. The evaluations page, and the API behind it, keeps the full judgment provenance so a supervisor reads *reason and authorship*, not a bare number.

## The ledger row: what a judgment carries

Every judgment is written as one `qa_evaluations` row in your tenant schema. Posting a score always goes through the Quality API; the server stamps the authorship fields, so a client can never claim to be a different judge than it is.

| Field         | What it records                                                                                                                            |
| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
| `form_id`     | Which rubric (evaluation form) was judged against.                                                                                         |
| `call_id`     | The call being graded.                                                                                                                     |
| `agent_id`    | The graded agent.                                                                                                                          |
| `reviewer_id` | The judge: a human user id, or a sentinel naming an automated judge.                                                                       |
| `scores`      | Per-criterion raw 0–100 scores; the weighted total is re-derived server-side from the form's weights, so no client can inflate the rollup. |
| `status`      | `pending` → `acknowledged` → (`appealed` → `resolved`) — the human reaction lifecycle.                                                     |

The derived field `evaluation_type` separates `self` from `reviewer` rows structurally: a row is `self` only when the evaluated agent authored it themselves (`agent_id === reviewer_id`).

### The judge registry: who writes `reviewer_id`

| Judge                                       | Trigger                                                                    | What it writes                                                                                                            |
| ------------------------------------------- | -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| Human reviewer (owner / admin / supervisor) | Manual assignment, or claiming a queued sample                             | A graded, reviewer-authored score that counts toward official averages.                                                   |
| LLM auto-scorer (`auto-llm`)                | Auto-scoring enabled in your tenant's QA settings                          | A machine-authored score on every completed, transcribed call; counted like a human score unless a reviewer overrides it. |
| Auto-sampler sentinel                       | A form with a sample share above zero queues a share of each agent's calls | An un-scored placeholder awaiting a human claim; excluded from averages until claimed.                                    |
| CSAT-trigger sentinel                       | A detractor CSAT response on an agent's call                               | An un-scored placeholder queued for human review; excluded until claimed.                                                 |
| Agent self-evaluation                       | The agent scoring their own call                                           | An immediately-`acknowledged` row excluded from official averages and the leaderboard.                                    |

## The LLM judge: trigger, thresholds, and correction path

Enable auto-scoring per tenant under your QA autoscore settings; it is **opt-in** — a tenant that never enables it sends zero transcripts to the LLM. When enabled, a background sweep finds completed calls with an extracted transcript and an answered agent, and scores each transcript against your active evaluation form under the head LLM judge.

Two thresholds shape what the judge writes:

* **Flag threshold** (default 70/100, configurable per tenant): an auto-score below the threshold marks the row for human review — the "needs human review" queue supervisors triage.
* **Flagging is triage, not removal.** A flagged row still counts toward averages; flagging routes it to a human for a second look.

The correction path is **override**: a reviewer replaces an LLM-authored row's scores with their own, and the weighted total is re-derived in place. Once overridden, a row reads as the reviewer's and exits the override queue. The self-vs-reviewer variance read pairs an agent's own self-evaluation with the official reviewer evaluation of the same call and returns the per-criterion delta — the target for coaching conversations.

## From score to downstream: coaching and leaderboards

Every ledger row routes to the same downstream family:

* **Low scores auto-trigger coaching.** A low QA score (or a low CSAT response) enrols the agent on a coaching plan with an open entry — no slack alert that gets lost.
* **The leaderboard and performance scorecards read the ledger** — they aggregate graded, non-self, non-placeholder rows, exactly as described in the [leaderboard concept](/concepts/qa-leaderboard-and-evaluations).
* **Coverage and readiness close the loop.** Autoscore coverage answers "how much of my traffic did the LLM judge actually score?"; auto-QA readiness answers "why not" when coverage sags — a disabled toggle, no active form, or calls without transcripts.

## The API shape

The payload the dashboard renders is the list endpoint's own shape; the page is a filter surface over it.

`GET /api/v1/quality/evaluations`

Query parameters (all optional):

| Parameter            | Filter                                                                                        |
| -------------------- | --------------------------------------------------------------------------------------------- |
| `agent_id`           | One agent.                                                                                    |
| `call_id`            | One call.                                                                                     |
| `reviewer_id`        | One judge (human id or sentinel).                                                             |
| `status`             | `pending` / `acknowledged` / `appealed` / `resolved`.                                         |
| `evaluation_type`    | `self` / `reviewer`.                                                                          |
| `auto_scored`        | `true` → rows authored by the LLM judge.                                                      |
| `flagged_for_review` | `true` → auto-score under your flag threshold.                                                |
| `csat_triggered`     | `true` → rows queued by the CSAT detractor trigger.                                           |
| `limit`              | Page size, 1–100 (default 50).                                                                |
| `cursor`             | Opaque keyset cursor echoed from `next_cursor`; malformed values return `400 INVALID_CURSOR`. |

Pagination is **keyset**: the response returns `next_cursor`; pass it back verbatim for the next page. Keyset paging stays O(rows returned) at any depth — important because the LLM judge appends one row per sampled call and the ledger grows continuously.

Role scoping: reviewers (owner, admin, supervisor) list across the org; an agent reads their own rows through the self-scoped guard. Related routes on the same journal:

| Operation                                               | Endpoint                                                                                  |
| ------------------------------------------------------- | ----------------------------------------------------------------------------------------- |
| List evaluations                                        | `GET /api/v1/quality/evaluations`                                                         |
| Self-vs-reviewer variance                               | `GET /api/v1/quality/evaluations/variance`                                                |
| Acknowledge / appeal / resolve / claim-score / override | `POST /api/v1/quality/evaluations/{id}/{acknowledge,appeal,resolve,claim-score,override}` |

The full request/response shapes are in the [Quality API reference](/api-reference/endpoints/quality), and the guide walks the dashboard triage loop in [Quality evaluations guide](/guides/quality-evaluations).

## Siblings in the loop

* [The call-quality evaluation lifecycle](/concepts/quality-evaluation-lifecycle) — assignments, calibration, coaching cards, coverage.
* [QA evaluations and the performance leaderboard](/concepts/qa-leaderboard-and-evaluations) — forms, scorecards, the application of ledger rows to rankings.
* [Quality evaluations guide](/guides/quality-evaluations) — the daily supervisor loop over the page.
