Skip to main content

Real-time agent-assist whisper coaching in the softphone

The browser softphone is where most agents actually take a call, so coaching lives there too — not on a separate call-detail page. While a call is in progress, an AI copilot reads the live transcript and whispers suggestions into the softphone’s active-call panel: a suggested reply, the next best action, up to three knowledge-base articles pulled from your workspace’s real KB, and a rolling sentiment read on the caller. Base path: /api/v1/voice Base surface: dashboard softphone pop-up or any SSE client; the card wire format is identical for both.

Where the panel lives

Open the softphone’s active-call view and toggle Agent assist. The toggle is persistent — turn it on once and the panel appears on every call until you turn it off; it never flashes on for calls you have not opened, because the backend issues cards only for calls the requester can already read. Each card carries:
  • suggested_response — one short paragraph the agent could say next, written to be spoken aloud (no greetings or sign-offs; the call is in progress).
  • kb_articles — up to three articles retrieved from your tenant’s own knowledge base. The model only distils the caller’s current informational need into a search query; the surfaced titles and snippets always come from real documents, each with document_id (and source_url where the source carries one) so the agent can open the article, never a hallucinated citation.
  • next_best_action — a single imperative next step (“offer a $20 credit”, “escalate to billing tier 2”).
  • suggested_action — where the next best step maps to a whitelisted one-click action (today: escalating to a ticket), a structured action object so the UI renders an execute button instead of plain text.
  • sentiment_score — a rolling [-1, +1] estimate of the caller’s current tone in the most recent turns, never an average over the whole call. Most ordinary conversations sit between −0.3 and +0.3, so treat −0.5 as genuinely unhappy, not merely quiet.
Cards arrive no more often than every eight seconds, and a card is emitted only when at least three new final transcript turns have landed since the previous card — the debounce on cadence. Only the trailing ~6,000 characters of the live transcript feed the model, so a long call costs a bounded amount per card.

Consuming the stream end-to-end over SSE

The softphone panel is one view of a single endpoint: GET /api/v1/voice/calls/{callId}/agent-assist/stream
The stream always opens with the same sequence, then carries cards until the call ends:
  1. event: connected — handshake with { "callId", "status" }.
  2. event: status — { "status", "terminal": true|false }. If the call is already in a terminal state (completed, failed, busy, no-answer, canceled), the stream sends event: end with { "reason": "terminal" } and closes cleanly — connect and subscribe, don’t poll for cards that will never come.
  3. event: assist — one frame per card. The data is the card JSON plus a server-issued envelope: monotonically increasing id, generated_at timestamp, and the model name.
  4. : ping — a comment-only keepalive line every 15 seconds so proxies and browsers see traffic even when no card is due.
  5. event: end — { "reason": "terminal" } when the call ends, or { "reason": "max-duration" } at the 30-minute hard cap; the client is expected to reconnect and continue from its last received id.
Here is a complete, runnable Node.js consumer that parses each of those frame types and defends against a stalling stream:
The reconnect helper is deliberately minimal — every reconnect gets a fresh connected/status pair, so resume from the highest assist-card id you have processed, not from an arbitrary frame. Access: the same voice-read scope as the call itself; agents only see cards for calls they can already open.

Tuning thresholds, debounce, and the transcript window

The coaching cadence is controlled by three numbers, all fixed and the same for every tenant:
  • Cadence gate — one card at most every 8 seconds and only after at least 3 new final transcript turns; a fast speaker produces a card roughly every three turns, a slow one never quicker than eight seconds. You tune cadence on the client side instead — render or discard, don’t re-fetch — so dashboards can show every fifth card or pin a favourite card without changing the stream.
  • Transcript window — the model sees the trailing ~6,000 characters (about a page and a half of speech) per card. The window never grows unbounded, so a 60-minute call costs the same per card as a 6-minute one. If your workflow needs long-range context (cross-call history, CRM records), fetch it in your own client and merge it with the cards — the endpoint intentionally carries only the rolling window.
  • Sentiment alarm threshold — the alert fires the moment the rolling score drops to −0.5 or below, the same boundary the post-call sentiment alerting uses. Keep per-agent per-call sensitivity on the client: e.g. suppress the supervisor nudge for an agent until the score stays below the threshold for two consecutive cards, or mark −0.8 as your own barge redline.

Building a custom coach widget

The softphone pop-up is one view; embedded phone bars and wall-panels can render the same cards from the same frames. A minimal render loop:
Two render rules matter:
  • When suggested_action is present, render a button. Today’s only whitelisted tool is escalate_to_ticket; press the button and the UI calls the existing ticket-creation endpoint with the bounded args.reason (max 300 characters), never a raw client-side tool invocation. When suggested_action is absent the next-best-action line stays plain text — never invent an execute button for unlisted actions.
  • Colour the sentiment strip by threshold: neutral above −0.5, warning at or below −0.5, your own team’s redline (often −0.8) below. The score is a rolling estimate of the most recent turns, so it can swing; smooth on the client with a short EMA if the strip flickers.

Supervisor lane: whisper vs barge

The sentiment emergency lane runs on the same cards the agent sees. When the rolling score drops to −0.5 or below, two things happen at once:
  • The supervisor wallboard gets a voice.sentiment.negative event with a recommended action: whisper (coach the agent silently) for a score of −0.7 or higher, barge (join audibly) below −0.7. Cadence is deliberately throttled: at most one alert per call per minute, so a persistent angry stretch lights the wallboard once, not once per card.
  • The agent’s own overlay gets an automated calming nudge instead of relying on the manual distress button.
Whisper is the right first move — the supervisor speaks to the agent without the caller hearing; barge is the escape hatch when the agent is drowning. The wallboard recommendation is advisory, not enforced; a supervisor who whispers at −0.9 has not broken anything, the recommendation just biases toward the safer intervention. The full listen / whisper / barge workflow and its audit ledger are covered in the supervisor live monitoring guide.

Failure modes and unhealthy streams

The pipeline is fail-soft at every junction — a degraded coaching layer never touches the call itself:
  • Card generation blows up — the call, transcript, and stream are untouched; the panel simply skips a tick. Diagnose by watching for a stream that opens connected and keeps : pinging but produces no assist frames for longer than a minute.
  • The call is already over when you connect — you get event: status with terminal: true, then event: end { "reason": "terminal" }. Reconnecting is pointless; use the post-call QA scorecard endpoints instead.
  • The stream goes silent — : ping comment keepalives run every 15 seconds, but the standard EventSource API suppresses comment lines, so a stalled socket shows up only as an onerror/close (or a transport-level parser) signal, not “no frames”. Treat that as a dead proxy or a pod restart — and rebuild — rather than as “no cards due”.
  • The 30-minute cap elapses — you get event: end { "reason": "max-duration" }; reconnect with the same URL, no resume token needed, and continue from the last received id.
  • Access wrong — cross-tenant or missing call ids return 404 before the stream opens; missing or insufficient credentials return 401/403 in the same envelope — test for the HTTP error before wiring the render loop.