Real-time agent-assist whisper coaching in the softphone
The browser softphone is where most agents actually take a call, so coaching lives there too — not on a separate call-detail page. While a call is in progress, an AI copilot reads the live transcript and whispers suggestions into the softphone’s active-call panel: a suggested reply, the next best action, up to three knowledge-base articles pulled from your workspace’s real KB, and a rolling sentiment read on the caller. Base path:/api/v1/voice
Base surface: dashboard softphone pop-up or any SSE client; the card wire format is identical for both.
Where the panel lives
Open the softphone’s active-call view and toggle Agent assist. The toggle is persistent — turn it on once and the panel appears on every call until you turn it off; it never flashes on for calls you have not opened, because the backend issues cards only for calls the requester can already read. Each card carries:suggested_response— one short paragraph the agent could say next, written to be spoken aloud (no greetings or sign-offs; the call is in progress).kb_articles— up to three articles retrieved from your tenant’s own knowledge base. The model only distils the caller’s current informational need into a search query; the surfaced titles and snippets always come from real documents, each withdocument_id(andsource_urlwhere the source carries one) so the agent can open the article, never a hallucinated citation.next_best_action— a single imperative next step (“offer a $20 credit”, “escalate to billing tier 2”).suggested_action— where the next best step maps to a whitelisted one-click action (today: escalating to a ticket), a structured action object so the UI renders an execute button instead of plain text.sentiment_score— a rolling [-1, +1] estimate of the caller’s current tone in the most recent turns, never an average over the whole call. Most ordinary conversations sit between −0.3 and +0.3, so treat −0.5 as genuinely unhappy, not merely quiet.
Consuming the stream end-to-end over SSE
The softphone panel is one view of a single endpoint:GET /api/v1/voice/calls/{callId}/agent-assist/stream
event: connected— handshake with{ "callId", "status" }.event: status—{ "status", "terminal": true|false }. If the call is already in a terminal state (completed,failed,busy,no-answer,canceled), the stream sendsevent: endwith{ "reason": "terminal" }and closes cleanly — connect and subscribe, don’t poll for cards that will never come.event: assist— one frame per card. Thedatais the card JSON plus a server-issued envelope: monotonically increasingid,generated_attimestamp, and themodelname.: ping— a comment-only keepalive line every 15 seconds so proxies and browsers see traffic even when no card is due.event: end—{ "reason": "terminal" }when the call ends, or{ "reason": "max-duration" }at the 30-minute hard cap; the client is expected to reconnect and continue from its last receivedid.
connected/status pair, so resume from the highest assist-card id you have processed, not from an arbitrary frame.
Access: the same voice-read scope as the call itself; agents only see cards for calls they can already open.
Tuning thresholds, debounce, and the transcript window
The coaching cadence is controlled by three numbers, all fixed and the same for every tenant:- Cadence gate — one card at most every 8 seconds and only after at least 3 new final transcript turns; a fast speaker produces a card roughly every three turns, a slow one never quicker than eight seconds. You tune cadence on the client side instead — render or discard, don’t re-fetch — so dashboards can show every fifth card or pin a favourite card without changing the stream.
- Transcript window — the model sees the trailing ~6,000 characters (about a page and a half of speech) per card. The window never grows unbounded, so a 60-minute call costs the same per card as a 6-minute one. If your workflow needs long-range context (cross-call history, CRM records), fetch it in your own client and merge it with the cards — the endpoint intentionally carries only the rolling window.
- Sentiment alarm threshold — the alert fires the moment the rolling score drops to −0.5 or below, the same boundary the post-call sentiment alerting uses. Keep per-agent per-call sensitivity on the client: e.g. suppress the supervisor nudge for an agent until the score stays below the threshold for two consecutive cards, or mark −0.8 as your own barge redline.
Building a custom coach widget
The softphone pop-up is one view; embedded phone bars and wall-panels can render the same cards from the same frames. A minimal render loop:- When
suggested_actionis present, render a button. Today’s only whitelisted tool isescalate_to_ticket; press the button and the UI calls the existing ticket-creation endpoint with the boundedargs.reason(max 300 characters), never a raw client-side tool invocation. Whensuggested_actionis absent the next-best-action line stays plain text — never invent an execute button for unlisted actions. - Colour the sentiment strip by threshold: neutral above −0.5, warning at or below −0.5, your own team’s redline (often −0.8) below. The score is a rolling estimate of the most recent turns, so it can swing; smooth on the client with a short EMA if the strip flickers.
Supervisor lane: whisper vs barge
The sentiment emergency lane runs on the same cards the agent sees. When the rolling score drops to −0.5 or below, two things happen at once:- The supervisor wallboard gets a
voice.sentiment.negativeevent with a recommended action: whisper (coach the agent silently) for a score of −0.7 or higher, barge (join audibly) below −0.7. Cadence is deliberately throttled: at most one alert per call per minute, so a persistent angry stretch lights the wallboard once, not once per card. - The agent’s own overlay gets an automated calming nudge instead of relying on the manual distress button.
Failure modes and unhealthy streams
The pipeline is fail-soft at every junction — a degraded coaching layer never touches the call itself:- Card generation blows up — the call, transcript, and stream are untouched; the panel simply skips a tick. Diagnose by watching for a stream that opens
connectedand keeps: pinging but produces noassistframes for longer than a minute. - The call is already over when you connect — you get
event: statuswithterminal: true, thenevent: end { "reason": "terminal" }. Reconnecting is pointless; use the post-call QA scorecard endpoints instead. - The stream goes silent —
: pingcomment keepalives run every 15 seconds, but the standardEventSourceAPI suppresses comment lines, so a stalled socket shows up only as anonerror/close(or a transport-level parser) signal, not “no frames”. Treat that as a dead proxy or a pod restart — and rebuild — rather than as “no cards due”. - The 30-minute cap elapses — you get
event: end { "reason": "max-duration" }; reconnect with the same URL, no resume token needed, and continue from the last receivedid. - Access wrong — cross-tenant or missing call ids return
404before the stream opens; missing or insufficient credentials return401/403in the same envelope — test for the HTTP error before wiring the render loop.
Related
- Supervisor live monitoring for voice calls — listen, whisper, barge, conference-join
- Call QA scorecards and auto-scoring — grade the same calls after the live coaching lane has done its work
- Live transcript SSE — agent-assist and supervisor-assist wire formats