> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orbit.devotel.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Live call transcription

> How real-time speech-to-text works on Classic and outbound calls, how provider failover and language routing behave, and where customer-supplied speech credentials apply.

# Live call transcription

Live call transcription turns speech on a voice call into text while the call is in progress. On Orbit, it runs on Classic voice calls and outbound voice calls when live transcription is enabled. The platform-managed speech chain can move between configured providers when a provider fails, subject to the failure rules below.

This is separate from three nearby features:

* The [STT Playground](/guides/stt-playground) compares speech vendors in a console. It is a test and comparison tool, not the live-call transcription path.
* AI voice agents have per-agent speech provider settings and a separate realtime media pipeline. See [Voice gateway: the realtime AI-voice media edge](/concepts/voice-gateway-realtime-edge).
* Post-call recording transcription runs after a recording is complete. See the [recording library](/voice/recording-library) for transcript search and recording workflows.

## Provider chain and language routing

The platform-managed chain is constructed once when the voice service starts, from providers whose credentials are provisioned. Deepgram is first, followed by Whisper, Azure Speech, and Google Speech in the fixed construction order. A provider without the required credentials is omitted, so the chain can contain fewer than four providers. The gateway describes this as: “The fallback chain is initialised once at server boot from whatever API keys are present — providers without credentials are simply omitted from the chain.”

For a requested language, Orbit reorders the configured chain so providers that declare support for that language are tried first. Providers that do not declare support move behind them but remain in the chain as fallbacks; language routing does not remove a provider. Region subtags such as `es-MX` use the primary language subtag (`es`). An unset language or `multi` keeps the original order.

You do not select a provider for an individual call. The provider order is established by the platform-managed chain and adjusted for the call's declared language; there is no per-call provider knob.

## What happens when a provider fails

The result depends on when the failure happens:

* **Before the first transcript:** Orbit tries the next configured provider and replays the buffered opening audio so the fallback can recognize speech that arrived before the switch. The buffer is capped at 4 MB. If it reaches the cap, replay is disabled for that call to bound memory use.
* **After transcripts, when the streaming provider exhausts its reconnect attempts:** Orbit switches to the next live-capable provider without replaying audio already consumed. Whisper is batch-only, so it is skipped for this live switch. Transcription resumes from new audio after the fallback connects; speech during the reconnect and switch window may not be transcribed.
* **Other errors after transcripts:** Orbit keeps the existing behavior and surfaces the error. It does not switch providers or replay audio, which could duplicate transcript text.
* **One-shot transcription:** For a complete audio buffer submitted for transcription, Orbit tries the language-ordered provider chain until one succeeds or the chain is exhausted. If every provider fails, the final provider error is returned.

A fallback can only help if another suitable provider is present in the chain. If the chain is exhausted, the error is surfaced to the call's transcription consumer; the voice call itself is not described as ending by this transcription behavior.

## Customer-supplied speech credentials

If you supply your own speech credential, your transcription stays within that credential's existing route boundary: you remain in control of your own transcription. The platform-managed failover chain described on this page does not take over that route, add its providers to it, or switch the call to a platform-managed credential. Your own provider's existing failure and recovery behavior continues to apply.

This boundary is distinct from choosing among platform-managed providers: tenant-supplied keys do not change the platform chain's boot-time provider availability or its order. For details on your configured speech credentials, see **Settings → Voice**. The AI-agent provider picker is also separate and continues to govern its own agent pipeline.

## Provisioning and operations

To ensure the platform-managed chain has fallback options, ask Devotel support to provision the platform speech providers needed for your deployment. The chain includes only providers available when the voice service starts; newly provisioned credentials are not added to an already-running chain until the service is restarted as part of the platform's rollout.

Use the [STT Playground](/guides/stt-playground) to check the catalogued availability of speech providers. For live mid-call recovery, verify that a streaming-capable provider is available after the primary in the configured chain; Whisper alone is not a live streaming backup. The Playground is an availability and comparison surface, not a per-call routing control.

During an outage, a failure before the first transcript may move to the next configured provider and replay up to 4 MB of opening audio. If a live stream exhausts reconnect attempts later in the call, the transcript may pause while Orbit connects the next live provider, then resume without replaying prior audio. Other post-transcript errors surface without a switch. If no eligible fallback is available, transcription cannot continue on another provider. For symptoms and recovery steps, see [Troubleshooting transcription and synthetic-voice failures](/troubleshooting/transcription-and-synthetic-voice-failures).

## Related reading

* [Voice gateway: the realtime AI-voice media edge](/concepts/voice-gateway-realtime-edge) — AI-agent speech-provider behavior.
* [STT Playground](/guides/stt-playground) — compare providers and check catalogued availability.
* [Troubleshooting transcription and synthetic-voice failures](/troubleshooting/transcription-and-synthetic-voice-failures) — diagnose speech-provider failures.
* [Live-call speech failover](/voice/live-call-stt-failover) — details on the reconnect-exhaustion switch for Classic and outbound calls.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.