Skip to main content

Wallboards and SLA telemetry

A supervisor board answers three different questions — how much work is in flight, how fast answers are flowing, and whether a service level is still being met — and Orbit drives all three from one model that spans voice queues and digital queues. Every surface on the wallboard is either a raw metric stream (the telemetry feeds) or a persisted policy object (the alarms and SLA rules the tenant poses over that stream). This page names that split once so the per-surface guides can stay thin. Inbound queue routing decides where a conversation lands; wallboards and SLA telemetry decide how knowledge of it becomes an operator intervention — an alarm, a page, a reroute, or an escalation. Everything in this model is tenant-owned governance over inbound queues; outbound (MT) voice and SMS exit only through the Devotel wholesale softswitch, and none of the wallboard/SLA surfaces touch any provider-side signalling path.

1. Metric taxonomy — the telemetry the board reads

All supervisor metrics reduce to one of four families, which the overview endpoint GET /api/v1/voice/supervisor/wallboard-hub aggregates in a single poll rather than from per-panel fetches: The voice live state (staffing + waiting) comes from the per-queue live snapshot a scheduler refreshes every minute; the volume/speed/SL rollup is aggregated over the last 24h of the queue-analytics interval buckets. The omnichannel board (GET /api/v1/voice/supervisor/omnichannel-wallboard) lifts this out of voice-only and joins the digital channels against it — per-channel live/unassigned/awaiting counts for chat, email, SMS, and the rest — so a blended floor reads one board instead of flipping between the inbox and the voice wallboard. The service-level formula is the one the analytics page and the hub must agree on exactly: offered minus short-abandons over answered-within-threshold (a caller who hangs up within a few seconds is excluded from the service-level denominator, which is the Five9/Genesys/NICE convention), and AHT = (talk + hold + wrap) / answered — queue wait time is not inflated into handle time.

2. Threshold alarm rules — a raw metric against a number

An alarm rule is the simplest policy object over one of the above metrics: a supervisor persists a row that says longest_wait > 120 for 60 seconds, and the surface fires when the stream keeps breaching it. The rule engine compares one metric against one threshold, and a raw alert is any of: a SSE frame, an in-app notification, a webhook, an email to the supervisor group. The evaluator that feeds it never decides what a breach means; it only declares that a metric crossed a threshold.
  • List — GET /api/v1/voice/wallboard/alarm-rules (optional queue_id, limit, offset; queue-scoped rules plus the tenant-wide rules, exactly the set the wallboard SLA scheduler evaluates for that queue each tick).
  • Create — POST /api/v1/voice/wallboard/alarm-rules. A rule carries queue_id (null/omitted = applies to every queue), name, metric, comparator, threshold, duration_seconds (0 = fire on the first breaching tick), channel (sse, sse+notification, webhook, email), and enabled. Creating an identical shape returns the existing row (200), so a double-click never creates duplicates.
  • Update / Delete — PATCH / DELETE /api/v1/voice/wallboard/alarm-rules/{ruleId}.
A duration_seconds window is the “sustained breach” guard — without it a one-tick spike pages the supervisor; with it the alarm only fires when the metric has been breaching for the whole window. Choose it with the alarm channel in mind: an email to the group can absorb a short window; a webhook to a paging surface should not fan out on noise. An alarm rule is the raw threshold half of the model. Distinguish it from the SLA policy below, which wraps the same telemetry in a governance object with escalation semantics — the former tells the supervisor that a metric crossed; the latter decides what happens next.

3. SLA policy — forecast vs breach-actual

A SLA policy is a persisted governance object the tenant stores under its own account, not a threshold a client passes per request. It answers “80% answered within 30 seconds, over a 15-minute window — if that breaks, page the supervisor, then open a ticket, and route the rest to the escalation ladder”. For voice queues the pair GET / PUT /api/v1/voice/queues/{queueId}/sla-breach/policy stores it, GET .../sla-breach/events reads the event log, and POST .../sla-breach/scan evaluates a measured snapshot against it. Digital queues have the parallel shape at GET / PUT /api/v1/inbox/digital-queues/{id}/sla-escalation/policy. The response branch of the policy (breachAction or escalationSteps) carries the remediation verb:
  • A single breachAction (open_ticket, page_supervisor, reroute_to_overflow_queue, enqueue_callback, none; digital queues use the smaller set notify_supervisors / page_supervisor / both / none).
  • A escalationSteps ladder — the ordered sequence of notify, reassign, slack_alert, webhook, slack_webhook, teams_webhook actions, each with a delayMinutesAfterBreach, that supersedes the single action whenever the ladder is not empty.
  • A breachCooldownSeconds dedupe rule so one breach does not page the supervisor every tick of the scheduler.
The forecast surface is the “SLA breach incoming” half of the model, sitting side-by-side against the actual-breach half: POST /api/v1/voice/queues/{queueId}/sla-forecast takes a measured window snapshot — targetPercentage, thresholdSeconds, evaluationWindowMinutes, offeredCalls, answeredWithinSla, queueDepth, oldestWaitSeconds, activeAgents, avgHandleTimeSeconds — and returns the foreseen breach ETA plus the agent-equivalent callback-capacity suggestion before the breach lands. The forecast uses the same service-level denominator the actual-breach path computes (offered minus short-abandons), so a supervisor reading a risky forecast against the same queue’s actual breach policy never sees two SL definitions disagree. The semantics to keep separate: The distinction matters operationally: the forecast pre-suggests callback capacity (so the queue can absorb a burst before the breach), while the breach path fires escalation (so the supervisor learns and re-routes after the fact). Both read the same service-level formula; a policy defines the target, the forecast merely projects it.

4. Digital queues — waiting-set telemetry and capacity reservations

For digital queues the telemetry is the waiting set each queue carries (the routing engine’s own FIFO backlog), not a telephony status:
  • Live stats per queue — GET /api/v1/inbox/digital-queues/stats (tenant-wide) and GET /api/v1/inbox/digital-queues/{queueId}/stats for one queue. Returns depth, longest-wait, available/busy headcount, and the queue’s overflow counters — the same numbers the routing engine takes its decisions against, served read-only so the numbers always reconcile with the engine’s own state.
  • Routing decision audit — GET /api/v1/inbox/digital-queues/{queueId}/decisions and GET /api/v1/inbox/routing-decisions/{conversationId} — “why did this conversation land where it landed”. Supervisors debug distribution on these, not on raw member availability.
  • Per-queue SLA escalation — GET / PUT / DELETE /api/v1/inbox/digital-queues/{id}/sla-escalation/policy plus GET .../events and POST .../scan: the same persisted object the voice sla-breach side defines, with an inbox-specific two-breach criterion (breached-conversation threshold and oldest-wait gate).
The omnichannel board’s digital tiles share the same scoping rules as the inbox SL attainment model: a first-response breach only counts when an inbound message actually arrived — an agent-initiated thread the customer never replied into cannot false-breach — and a closed / resolved / archived thread is never live work. The live voice state and the digital tiles are joined into one supervisor surface by GET /api/v1/voice/supervisor/omnichannel-wallboard. Omnichannel-capacity reservations. The blended loadScore is the read-only signal that voice-queue routing, digital-queue routing, and the omnichannel work-item router all share: an agent available on voice but saturated on open digital conversations is skipped by the load score before the queue’s hard state gate (available) gets consulted. The capacity stream is GET /api/v1/voice/capacity/events (SSE), which the supervisor board subscribes to; a bare /api/v1/voice/capacity hit is permanently redirected into the canonical path. The capacity picture is a suggestion the dispatcher reads; the presence gate remains the rule, exactly as the ACD queue model defines it.

5. Where the analytics sources live

Telemetry is derived from source-of-truth records at read time, never from a separate metrics store on the request path:
  • Voice queue analytics — GET /api/v1/voice/queues/{id}/analytics?from=&to=&interval= returns the windowed rollup (offered/answered/abandoned/SL/AHT/ASA/abandon/occupancy) plus the First-Call-Resolution heuristic, aggregated in the same buckets the wallboard hub sums.
  • Live SSE stream — GET /api/v1/voice/queues/{id}/live — emits per-tick snapshots to the supervisor board instead of rescanning the interval buckets on every poll.
  • Digital-queue live stats — the per-queue waiting set from the digital-queue stats routes above.
  • Hub aggregates — GET /api/v1/voice/supervisor/wallboard-hub for “the board in one payload” (SLA, live, quality); the route argument in the register — consistency across polling surfaces — was the reason it lands in one round trip.
For how queue truth reconciles monthly (the SLA attestation and the Reliability report), see the SLA attestation model. For the queue model each of these metrics is drawn from, see the ACD queue model. Supervisor-level live monitoring live on the voice queues guide and the omnichannel queue routing guide.