> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Wallboards and SLA telemetry: metric families, threshold alarms, and SLA policy

> How Orbit models the supervisor board — the queue metric taxonomy (depth, AHT, ASA, occupancy, service level), how raw threshold alarms differ from persisted SLA breach policies, the forecast-vs-actual SLA semantics, and the digital-queue surfaces that share the same model.

# Wallboards and SLA telemetry

A supervisor board answers three different questions — **how much work is in flight**, **how fast answers are flowing**, and **whether a service level is still being met** — and Orbit drives all three from one model that spans voice queues and digital queues. Every surface on the wallboard is either a **raw metric stream** (the telemetry feeds) or a **persisted policy object** (the alarms and SLA rules the tenant poses over that stream). This page names that split once so the per-surface guides can stay thin.

Inbound queue routing decides **where** a conversation lands; wallboards and SLA telemetry decide **how knowledge of it becomes an operator intervention** — an alarm, a page, a reroute, or an escalation. Everything in this model is tenant-owned governance over inbound queues; outbound (MT) voice and SMS exit only through the Devotel wholesale softswitch, and none of the wallboard/SLA surfaces touch any provider-side signalling path.

## 1. Metric taxonomy — the telemetry the board reads

All supervisor metrics reduce to one of four families, which the overview endpoint `GET /api/v1/voice/supervisor/wallboard-hub` aggregates in a single poll rather than from per-panel fetches:

| Family            | The metrics in it                                                      | What it answers                                            |
| ----------------- | ---------------------------------------------------------------------- | ---------------------------------------------------------- |
| **Volume**        | `offered`, `answered`, `abandoned`                                     | How much work arrived, and what happened to it             |
| **Speed**         | AHT (average handle time), ASA (average speed of answer), longest-wait | How fast answers are flowing                               |
| **Staffing**      | agents logged in / available / on call / in wrap; occupancy            | Whether the floor has enough head to take the next arrival |
| **Service level** | the SLA % against the queue's target                                   | The managed objective being tracked                        |

The voice live state (staffing + waiting) comes from the per-queue live snapshot a scheduler refreshes every minute; the volume/speed/SL rollup is aggregated over the last 24h of the queue-analytics interval buckets. The omnichannel board (`GET /api/v1/voice/supervisor/omnichannel-wallboard`) lifts this out of voice-only and joins the digital channels against it — per-channel live/unassigned/awaiting counts for chat, email, SMS, and the rest — so a blended floor reads one board instead of flipping between the inbox and the voice wallboard.

The service-level formula is the one the analytics page and the hub must agree on exactly: **offered minus short-abandons over answered-within-threshold** (a caller who hangs up within a few seconds is excluded from the service-level denominator, which is the Five9/Genesys/NICE convention), and AHT = (talk + hold + wrap) / answered — queue wait time is not inflated into handle time.

## 2. Threshold alarm rules — a raw metric against a number

An alarm rule is the simplest policy object over one of the above metrics: a supervisor persists a row that says `longest_wait > 120` for 60 seconds, and the surface fires when the stream keeps breaching it. The rule engine compares one metric against one threshold, and a raw alert is any of: a SSE frame, an in-app notification, a webhook, an email to the supervisor group. The evaluator that feeds it never decides what a breach *means*; it only declares that a metric crossed a threshold.

```bash theme={null}
curl "https://api.orbit.devotel.io/api/v1/voice/wallboard/alarm-rules" \
  -H "X-API-Key: $ORBIT_API_KEY"
```

* **List** — `GET /api/v1/voice/wallboard/alarm-rules` (optional `queue_id`, `limit`, `offset`; queue-scoped rules plus the tenant-wide rules, exactly the set the wallboard SLA scheduler evaluates for that queue each tick).
* **Create** — `POST /api/v1/voice/wallboard/alarm-rules`. A rule carries `queue_id` (null/omitted = applies to every queue), `name`, `metric`, `comparator`, `threshold`, `duration_seconds` (0 = fire on the first breaching tick), `channel` (`sse`, `sse+notification`, `webhook`, `email`), and `enabled`. Creating an identical shape returns the existing row (200), so a double-click never creates duplicates.
* **Update / Delete** — `PATCH / DELETE /api/v1/voice/wallboard/alarm-rules/{ruleId}`.

A `duration_seconds` window is the "sustained breach" guard — without it a one-tick spike pages the supervisor; with it the alarm only fires when the metric has been breaching for the whole window. Choose it with the alarm channel in mind: an email to the group can absorb a short window; a webhook to a paging surface should not fan out on noise.

An alarm rule is the **raw threshold** half of the model. Distinguish it from the SLA policy below, which wraps the same telemetry in a governance object with escalation semantics — the former tells the supervisor *that* a metric crossed; the latter decides *what happens next*.

## 3. SLA policy — forecast vs breach-actual

A SLA policy is a **persisted governance object the tenant stores under its own account**, not a threshold a client passes per request. It answers "80% answered within 30 seconds, over a 15-minute window — if that breaks, page the supervisor, then open a ticket, and route the rest to the escalation ladder". For voice queues the pair `GET / PUT /api/v1/voice/queues/{queueId}/sla-breach/policy` stores it, `GET .../sla-breach/events` reads the event log, and `POST .../sla-breach/scan` evaluates a measured snapshot against it. Digital queues have the parallel shape at `GET / PUT /api/v1/inbox/digital-queues/{id}/sla-escalation/policy`.

The response branch of the policy (`breachAction` or `escalationSteps`) carries the remediation verb:

* A single `breachAction` (`open_ticket`, `page_supervisor`, `reroute_to_overflow_queue`, `enqueue_callback`, `none`; digital queues use the smaller set `notify_supervisors` / `page_supervisor` / `both` / `none`).
* A `escalationSteps` ladder — the ordered sequence of `notify`, `reassign`, `slack_alert`, `webhook`, `slack_webhook`, `teams_webhook` actions, each with a `delayMinutesAfterBreach`, that supersedes the single action whenever the ladder is not empty.
* A `breachCooldownSeconds` dedupe rule so one breach does not page the supervisor every tick of the scheduler.

The forecast surface is the "SLA breach *incoming*" half of the model, sitting side-by-side against the actual-breach half: `POST /api/v1/voice/queues/{queueId}/sla-forecast` takes a measured window snapshot — `targetPercentage`, `thresholdSeconds`, `evaluationWindowMinutes`, `offeredCalls`, `answeredWithinSla`, `queueDepth`, `oldestWaitSeconds`, `activeAgents`, `avgHandleTimeSeconds` — and returns the foreseen breach ETA plus the agent-equivalent callback-capacity suggestion before the breach lands. The forecast uses the same service-level denominator the actual-breach path computes (offered minus short-abandons), so a supervisor reading a risky forecast against the same queue's actual breach policy never sees two SL definitions disagree.

The semantics to keep separate:

| Surface                                               | What it measures                                                                                                             | When the answer changes                                                         |
| ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- |
| **Forecast** (`POST .../sla-forecast`)                | *Suppose* the queue keeps draining at the measured AHT/headcount; how many minutes until the oldest backlog breach would hit | Changes on every live tick; use it to act before the breach                     |
| **Breach actual** (`.../sla-breach/scan` + event log) | *Actually* how the queue scored against its stored SLA policy over the rolling window                                        | Changes only when the window's observed service level crosses the stored target |

The distinction matters operationally: the forecast pre-suggests callback capacity (so the queue can absorb a burst before the breach), while the breach path fires escalation (so the supervisor learns and re-routes after the fact). Both read the same service-level formula; a policy defines the target, the forecast merely projects it.

## 4. Digital queues — waiting-set telemetry and capacity reservations

For digital queues the telemetry is the **waiting set** each queue carries (the routing engine's own FIFO backlog), not a telephony status:

* **Live stats per queue** — `GET /api/v1/inbox/digital-queues/stats` (tenant-wide) and `GET /api/v1/inbox/digital-queues/{queueId}/stats` for one queue. Returns depth, longest-wait, available/busy headcount, and the queue's overflow counters — the same numbers the routing engine takes its decisions against, served read-only so the numbers always reconcile with the engine's own state.
* **Routing decision audit** — `GET /api/v1/inbox/digital-queues/{queueId}/decisions` and `GET /api/v1/inbox/routing-decisions/{conversationId}` — "why did this conversation land where it landed". Supervisors debug distribution on these, not on raw member availability.
* **Per-queue SLA escalation** — `GET / PUT / DELETE /api/v1/inbox/digital-queues/{id}/sla-escalation/policy` plus `GET .../events` and `POST .../scan`: the same persisted object the voice `sla-breach` side defines, with an inbox-specific two-breach criterion (breached-conversation threshold and oldest-wait gate).

The omnichannel board's digital tiles share the same scoping rules as the inbox SL attainment model: a first-response breach only counts when an inbound message actually arrived — an agent-initiated thread the customer never replied into cannot false-breach — and a closed / resolved / archived thread is never live work. The live voice state and the digital tiles are joined into one supervisor surface by `GET /api/v1/voice/supervisor/omnichannel-wallboard`.

**Omnichannel-capacity reservations.** The blended `loadScore` is the read-only signal that voice-queue routing, digital-queue routing, and the omnichannel work-item router all share: an agent `available` on voice but saturated on open digital conversations is skipped by the load score before the queue's hard state gate (`available`) gets consulted. The capacity stream is `GET /api/v1/voice/capacity/events` (SSE), which the supervisor board subscribes to; a bare `/api/v1/voice/capacity` hit is permanently redirected into the canonical path. The capacity picture is a suggestion the dispatcher reads; the presence gate remains the rule, exactly as the ACD queue model defines it.

## 5. Where the analytics sources live

Telemetry is derived from source-of-truth records at read time, never from a separate metrics store on the request path:

* **Voice queue analytics** — `GET /api/v1/voice/queues/{id}/analytics?from=&to=&interval=` returns the windowed rollup (offered/answered/abandoned/SL/AHT/ASA/abandon/occupancy) plus the First-Call-Resolution heuristic, aggregated in the same buckets the wallboard hub sums.
* **Live SSE stream** — `GET /api/v1/voice/queues/{id}/live` — emits per-tick snapshots to the supervisor board instead of rescanning the interval buckets on every poll.
* **Digital-queue live stats** — the per-queue waiting set from the digital-queue stats routes above.
* **Hub aggregates** — `GET /api/v1/voice/supervisor/wallboard-hub` for "the board in one payload" (SLA, live, quality); the route argument in the register — consistency across polling surfaces — was the reason it lands in one round trip.

For how queue truth reconciles monthly (the SLA attestation and the Reliability report), see the [SLA attestation model](/concepts/sla-attestation-model). For the queue model each of these metrics is drawn from, see the [ACD queue model](/concepts/acd-queue-model). Supervisor-level live monitoring live on the [voice queues guide](/guides/voice-queues) and the omnichannel queue routing guide.
