Live-call speech failover
A primary speech vendor’s live connection can die mid-call: a vendor edge node restarts, a TLS session drops and refuses to come back, a regional outage silences the stream. Before 2026-10-11, once a live session had delivered its first transcript, a dead primary connection ended live transcription for that call: the error surfaced, transcription stopped, and the rest of the conversation went untranscribed. Now the live transcription layer recovers in that one specific failure mode: when the primary streaming connection exhausts its own reconnection budget — every retry the vendor adapter was willing to make has been made, and the connection is permanently dead — the call moves to your next configured live speech vendor and transcription continues.Which calls fail over
Two call surfaces take the failover:- Classic voice calls — inbound calls answered through the platform’s call flows (queues, ring groups, IVR, forwarding) with live transcription enabled.
- Outbound voice calls — calls the platform originates (dialer and direct outbound) with live transcription enabled.
- AI voice agents. Each agent’s speech vendor is the one you picked in the per-agent picker, governed by the agent pipeline’s own provider contract (see Voice gateway: the realtime AI-voice media edge, section 3). The agent surface does not take the mid-call switch.
- Customer-supplied speech credentials. If you wire your own speech key (BYO), your transcription runs inside your credential’s route boundary and keeps its existing single-vendor behavior. The failover described here applies to platform-managed connections only.
The switching rule
When the failover fires, the call advances to the next live vendor in the platform’s provider chain, skipping Whisper:- The platform STT chain is constructed once at service boot from the vendors with provisioned credentials. Its order is fixed by that construction — there is no per-tenant or per-call provider ordering knob. A call inherits whichever chain the platform booted with.
- Whisper is skipped for the live path by design. Whisper is a batch-only speech model: it transcribes whole audio segments after they are complete, not a live stream. A live call cannot be handed to a batch recognizer mid-utterance, so the switch walks past Whisper to the next vendor that accepts a live stream.
- The call does not re-rank or re-sort: it takes the next live-capable vendor below the failed one in chain order. If no live vendor below the failed one is provisioned, the call keeps the pre-2026-10-11 behavior and the transcription error surfaces instead.
Why there is no audio replay on an exhaustion switch
Some platform failovers replay the audio the caller already spoke so the opening of the utterance is not lost — the pre-transcript failover on the AI voice agent edge does exactly that, replaying a bounded preroll buffer into the starting vendor (again, see Voice gateway: the realtime AI-voice media edge, section 3). The exhaustion switch deliberately does not:- The switch can only fire after the primary connection has already delivered transcripts — exhaustion means the stream was live and transcribing, not that it failed to start. The audio spoken so far has already been sent to the primary and (up to the failure point) already transcribed.
- Replaying that same audio into the fallback would produce duplicate or out-of-order transcripts — the take-it-once pipeline would see the opening utterance twice, and any turn routing or summary built on the transcript would double-count it.
- So the fallback connection starts clean, from the moment it opens. Whatever the caller said while the primary was collapsing and reconnect attempts were burning — typically a couple of seconds of speech — is not re-read. Transcripts resume from the next words spoken.
What you observe as a tenant
There is nothing to configure and nothing to opt into. The behavior applies automatically to Classic and outbound calls running on platform-managed speech connections.- Calls without a live backup keep their existing behavior. If your platform chain has no live vendor below the primary — no secondary streaming vendor is provisioned — a reconnect-exhaustion ends live transcription the same way it did before, with the error surfaced. The switch only has somewhere to go if a backup exists.
- A failover is invisible except as continuity. The call’s consumers (live transcript dialers, summaries, recording annotations) see the transcript stream pause for the reconnect window and then resume on the backup. There is no special event a tenant integration must handle — transcription simply keeps flowing.
- Confirming you have a backup. Check which speech vendors are
provisioned for your workspace: the STT
Playground lists every catalogued vendor with an
Available/Not yet availablestatus, and Settings → Voice shows your BYO speech keys if you supply your own. A second platform-managed live vendor downstream of your primary is what the switch reaches for — if only one live vendor is provisioned, live calls have no failover destination. Contact Devotel support if you need a second streaming vendor enabled on your workspace. - Own keys stay own. If you supply your own speech credential, your calls never take the platform failover — they fail and recover inside your credential’s boundary, as before.
What did not change
- Mid-stream errors that are not reconnect-exhaustion still surface without a switch. A vendor connection that dies after one transcript for any other reason keeps the established surface-the-error behavior — the platform distinguishes “connection is permanently dead” from “connection hiccuped”, and only the first one triggers a live switch.
- TTS is untouched. The synthetic-voice side of a call has its own failover discipline and did not change.
- Billing shape is unchanged. Speech metering applies per-vendor as it always has; a call that crosses vendors mid-stream meters each leg on the vendor that served it.
Related reading
- Voice gateway: the realtime AI-voice media edge — the provider-chain contract for AI voice agents, including the pre-stream replay semantics this page contrasts with.
- STT Playground — compare speech vendors and confirm which are provisioned on your workspace.
- Troubleshooting: transcription, synthesis, and deepfake failures — the error-code matrix for the produce-then-read speech pipeline, including upstream vendor failures.