Skip to main content

The LLM-judged evaluation ledger

Read the call-quality evaluation lifecycle for the operations layer around scoring; read QA evaluations and the performance leaderboard for the evaluation form and leaderboard contract. This page names the thing both pipeline concepts leave implicit: the evaluation ledger — the per-call judgment record the Quality → Evaluations page renders, where every row carries who judged it (an LLM scorer, a human reviewer, or a queued placeholder), what triggered the judgment, and where the pass/fail decision goes next.

Why the judge matters, not just the score

A leaderboard compresses each agent to a QA average. When that average drops, the supervisor’s first question is never “how low” — it is “who said so, and which to trust.” Different kinds of judges produce different trust levels:
  • A row authored by a human reviewer is an official audit-grade judgment with an acknowledge/appeal lifecycle.
  • A row authored by the auto-scoring LLM judge is a machine judgment — correctable by a human override, but until overridden it counts toward the official average exactly like a human score.
  • A queued placeholder — an auto-sampling or CSAT-trigger row with a zero pending score — is not a judgment at all and is excluded from every average until a human claims and scores it.
  • A self-evaluation — the agent scoring their own call — is coaching material, deliberately excluded from official averages and the leaderboard so agents cannot inflate their own rank.
The average alone cannot answer any of that. The evaluations page, and the API behind it, keeps the full judgment provenance so a supervisor reads reason and authorship, not a bare number.

The ledger row: what a judgment carries

Every judgment is written as one qa_evaluations row in your tenant schema. Posting a score always goes through the Quality API; the server stamps the authorship fields, so a client can never claim to be a different judge than it is. The derived field evaluation_type separates self from reviewer rows structurally: a row is self only when the evaluated agent authored it themselves (agent_id === reviewer_id).

The judge registry: who writes reviewer_id

The LLM judge: trigger, thresholds, and correction path

Enable auto-scoring per tenant under your QA autoscore settings; it is opt-in — a tenant that never enables it sends zero transcripts to the LLM. When enabled, a background sweep finds completed calls with an extracted transcript and an answered agent, and scores each transcript against your active evaluation form under the head LLM judge. Two thresholds shape what the judge writes:
  • Flag threshold (default 70/100, configurable per tenant): an auto-score below the threshold marks the row for human review — the “needs human review” queue supervisors triage.
  • Flagging is triage, not removal. A flagged row still counts toward averages; flagging routes it to a human for a second look.
The correction path is override: a reviewer replaces an LLM-authored row’s scores with their own, and the weighted total is re-derived in place. Once overridden, a row reads as the reviewer’s and exits the override queue. The self-vs-reviewer variance read pairs an agent’s own self-evaluation with the official reviewer evaluation of the same call and returns the per-criterion delta — the target for coaching conversations.

From score to downstream: coaching and leaderboards

Every ledger row routes to the same downstream family:
  • Low scores auto-trigger coaching. A low QA score (or a low CSAT response) enrols the agent on a coaching plan with an open entry — no slack alert that gets lost.
  • The leaderboard and performance scorecards read the ledger — they aggregate graded, non-self, non-placeholder rows, exactly as described in the leaderboard concept.
  • Coverage and readiness close the loop. Autoscore coverage answers “how much of my traffic did the LLM judge actually score?”; auto-QA readiness answers “why not” when coverage sags — a disabled toggle, no active form, or calls without transcripts.

The API shape

The payload the dashboard renders is the list endpoint’s own shape; the page is a filter surface over it. GET /api/v1/quality/evaluations Query parameters (all optional): Pagination is keyset: the response returns next_cursor; pass it back verbatim for the next page. Keyset paging stays O(rows returned) at any depth — important because the LLM judge appends one row per sampled call and the ledger grows continuously. Role scoping: reviewers (owner, admin, supervisor) list across the org; an agent reads their own rows through the self-scoped guard. Related routes on the same journal: The full request/response shapes are in the Quality API reference, and the guide walks the dashboard triage loop in Quality evaluations guide.

Siblings in the loop