Wallboards and SLA telemetry
A supervisor board answers three different questions — how much work is in flight, how fast answers are flowing, and whether a service level is still being met — and Orbit drives all three from one model that spans voice queues and digital queues. Every surface on the wallboard is either a raw metric stream (the telemetry feeds) or a persisted policy object (the alarms and SLA rules the tenant poses over that stream). This page names that split once so the per-surface guides can stay thin. Inbound queue routing decides where a conversation lands; wallboards and SLA telemetry decide how knowledge of it becomes an operator intervention — an alarm, a page, a reroute, or an escalation. Everything in this model is tenant-owned governance over inbound queues; outbound (MT) voice and SMS exit only through the Devotel wholesale softswitch, and none of the wallboard/SLA surfaces touch any provider-side signalling path.1. Metric taxonomy — the telemetry the board reads
All supervisor metrics reduce to one of four families, which the overview endpointGET /api/v1/voice/supervisor/wallboard-hub aggregates in a single poll rather than from per-panel fetches:
The voice live state (staffing + waiting) comes from the per-queue live snapshot a scheduler refreshes every minute; the volume/speed/SL rollup is aggregated over the last 24h of the queue-analytics interval buckets. The omnichannel board (
GET /api/v1/voice/supervisor/omnichannel-wallboard) lifts this out of voice-only and joins the digital channels against it — per-channel live/unassigned/awaiting counts for chat, email, SMS, and the rest — so a blended floor reads one board instead of flipping between the inbox and the voice wallboard.
The service-level formula is the one the analytics page and the hub must agree on exactly: offered minus short-abandons over answered-within-threshold (a caller who hangs up within a few seconds is excluded from the service-level denominator, which is the Five9/Genesys/NICE convention), and AHT = (talk + hold + wrap) / answered — queue wait time is not inflated into handle time.
2. Threshold alarm rules — a raw metric against a number
An alarm rule is the simplest policy object over one of the above metrics: a supervisor persists a row that sayslongest_wait > 120 for 60 seconds, and the surface fires when the stream keeps breaching it. The rule engine compares one metric against one threshold, and a raw alert is any of: a SSE frame, an in-app notification, a webhook, an email to the supervisor group. The evaluator that feeds it never decides what a breach means; it only declares that a metric crossed a threshold.
- List —
GET /api/v1/voice/wallboard/alarm-rules(optionalqueue_id,limit,offset; queue-scoped rules plus the tenant-wide rules, exactly the set the wallboard SLA scheduler evaluates for that queue each tick). - Create —
POST /api/v1/voice/wallboard/alarm-rules. A rule carriesqueue_id(null/omitted = applies to every queue),name,metric,comparator,threshold,duration_seconds(0 = fire on the first breaching tick),channel(sse,sse+notification,webhook,email), andenabled. Creating an identical shape returns the existing row (200), so a double-click never creates duplicates. - Update / Delete —
PATCH / DELETE /api/v1/voice/wallboard/alarm-rules/{ruleId}.
duration_seconds window is the “sustained breach” guard — without it a one-tick spike pages the supervisor; with it the alarm only fires when the metric has been breaching for the whole window. Choose it with the alarm channel in mind: an email to the group can absorb a short window; a webhook to a paging surface should not fan out on noise.
An alarm rule is the raw threshold half of the model. Distinguish it from the SLA policy below, which wraps the same telemetry in a governance object with escalation semantics — the former tells the supervisor that a metric crossed; the latter decides what happens next.
3. SLA policy — forecast vs breach-actual
A SLA policy is a persisted governance object the tenant stores under its own account, not a threshold a client passes per request. It answers “80% answered within 30 seconds, over a 15-minute window — if that breaks, page the supervisor, then open a ticket, and route the rest to the escalation ladder”. For voice queues the pairGET / PUT /api/v1/voice/queues/{queueId}/sla-breach/policy stores it, GET .../sla-breach/events reads the event log, and POST .../sla-breach/scan evaluates a measured snapshot against it. Digital queues have the parallel shape at GET / PUT /api/v1/inbox/digital-queues/{id}/sla-escalation/policy.
The response branch of the policy (breachAction or escalationSteps) carries the remediation verb:
- A single
breachAction(open_ticket,page_supervisor,reroute_to_overflow_queue,enqueue_callback,none; digital queues use the smaller setnotify_supervisors/page_supervisor/both/none). - A
escalationStepsladder — the ordered sequence ofnotify,reassign,slack_alert,webhook,slack_webhook,teams_webhookactions, each with adelayMinutesAfterBreach, that supersedes the single action whenever the ladder is not empty. - A
breachCooldownSecondsdedupe rule so one breach does not page the supervisor every tick of the scheduler.
POST /api/v1/voice/queues/{queueId}/sla-forecast takes a measured window snapshot — targetPercentage, thresholdSeconds, evaluationWindowMinutes, offeredCalls, answeredWithinSla, queueDepth, oldestWaitSeconds, activeAgents, avgHandleTimeSeconds — and returns the foreseen breach ETA plus the agent-equivalent callback-capacity suggestion before the breach lands. The forecast uses the same service-level denominator the actual-breach path computes (offered minus short-abandons), so a supervisor reading a risky forecast against the same queue’s actual breach policy never sees two SL definitions disagree.
The semantics to keep separate:
The distinction matters operationally: the forecast pre-suggests callback capacity (so the queue can absorb a burst before the breach), while the breach path fires escalation (so the supervisor learns and re-routes after the fact). Both read the same service-level formula; a policy defines the target, the forecast merely projects it.
4. Digital queues — waiting-set telemetry and capacity reservations
For digital queues the telemetry is the waiting set each queue carries (the routing engine’s own FIFO backlog), not a telephony status:- Live stats per queue —
GET /api/v1/inbox/digital-queues/stats(tenant-wide) andGET /api/v1/inbox/digital-queues/{queueId}/statsfor one queue. Returns depth, longest-wait, available/busy headcount, and the queue’s overflow counters — the same numbers the routing engine takes its decisions against, served read-only so the numbers always reconcile with the engine’s own state. - Routing decision audit —
GET /api/v1/inbox/digital-queues/{queueId}/decisionsandGET /api/v1/inbox/routing-decisions/{conversationId}— “why did this conversation land where it landed”. Supervisors debug distribution on these, not on raw member availability. - Per-queue SLA escalation —
GET / PUT / DELETE /api/v1/inbox/digital-queues/{id}/sla-escalation/policyplusGET .../eventsandPOST .../scan: the same persisted object the voicesla-breachside defines, with an inbox-specific two-breach criterion (breached-conversation threshold and oldest-wait gate).
GET /api/v1/voice/supervisor/omnichannel-wallboard.
Omnichannel-capacity reservations. The blended loadScore is the read-only signal that voice-queue routing, digital-queue routing, and the omnichannel work-item router all share: an agent available on voice but saturated on open digital conversations is skipped by the load score before the queue’s hard state gate (available) gets consulted. The capacity stream is GET /api/v1/voice/capacity/events (SSE), which the supervisor board subscribes to; a bare /api/v1/voice/capacity hit is permanently redirected into the canonical path. The capacity picture is a suggestion the dispatcher reads; the presence gate remains the rule, exactly as the ACD queue model defines it.
5. Where the analytics sources live
Telemetry is derived from source-of-truth records at read time, never from a separate metrics store on the request path:- Voice queue analytics —
GET /api/v1/voice/queues/{id}/analytics?from=&to=&interval=returns the windowed rollup (offered/answered/abandoned/SL/AHT/ASA/abandon/occupancy) plus the First-Call-Resolution heuristic, aggregated in the same buckets the wallboard hub sums. - Live SSE stream —
GET /api/v1/voice/queues/{id}/live— emits per-tick snapshots to the supervisor board instead of rescanning the interval buckets on every poll. - Digital-queue live stats — the per-queue waiting set from the digital-queue stats routes above.
- Hub aggregates —
GET /api/v1/voice/supervisor/wallboard-hubfor “the board in one payload” (SLA, live, quality); the route argument in the register — consistency across polling surfaces — was the reason it lands in one round trip.