STT Playground
Picking a speech-to-text vendor for a voice agent is a transcription-quality decision, and it is hard to make blind. The STT Playground lets you try a short clip against every speech-to-text vendor Orbit recognises — from the dashboard, with no API integration — and read each vendor’s transcript, confidence, and latency side by side before you commit one to an agent. Open it under Voice → STT Playground in the dashboard. It is gated to theowner, admin, and developer roles because each comparison run invokes
metered upstream transcription calls. If a BYO STT key is rejected while you
configure voice inference, see Troubleshooting: voice inference credential
rejected.
What it compares
Each run sends your clip to every catalogued STT vendor independently and returns one row per vendor:
Vendors are compared in parallel. If one vendor fails — bad credentials, an
upstream outage — it gets its own error row and the rest of the comparison
still comes back. One bad vendor never sinks the run.
Honest availability
Today, Deepgram is the vendor with a provisioned batch transcription path. The other catalogued vendors — Soniox, Speechmatics, AssemblyAI, and xAI — are recognised by Orbit and appear in the comparison, but are not yet provisioned. They show as “Not yet available” rather than a fabricated result. That honesty is deliberate: the playground shows you the same availability the per-agent STT picker sees, so the comparison never promises a vendor you cannot actually wire. A vendor’s status flips toAvailable on its own once its
upstream credentials are provisioned.
Record or upload
Two ways to get a clip in — both feed the same comparison:- Record from your microphone. Use the mic button on the page; speak the kind of prompt a caller would actually say (numbers, names, accents, background noise — whatever your agents will face).
- Upload a clip. Accepted formats are MP3, WAV, OGG, AAC, and M4A, up to 25 MB. Keep clips short and representative: one prompt, not a whole call.
Upload contract
The page calls a single endpoint, and the same contract applies if you script your own comparison:POST /api/v1/voice/stt-preview
multipart/form-datawith onefilefield holding the audio clip.- Optional
vendorsform field: a comma-separated list of vendor ids to restrict the run (for exampledeepgram). Omit it to compare every catalogued vendor. - Auth: an API key or session token with the
voice:writescope, on anowner,admin, ordeveloperrole. - Response:
{ results: [{ vendor, label, provisioned, transcript?, confidence?, latency_ms?, error? }] }— one entry per vendor compared.
415 when the upload is not multipart or not an accepted
audio type, 422 when no file (or an empty file) was provided, and 413 when
the clip exceeds the 25 MB limit.
Script your own comparison
The playground page is the samePOST /api/v1/voice/stt-preview endpoint you
can call yourself, so a comparison run drops straight into a benchmark script
or a CI check that re-tests your reference clips.
Upload a clip with curl — one file form field, and an optional vendors
field to narrow the run:
--form vendors=... to compare every catalogued vendor.
A 200 response returns one entry per vendor compared. An unprovisioned vendor
comes back as its own provisioned: false entry rather than failing the run,
and an upstream failure lands on that vendor’s error field alone — one bad
vendor never sinks the rest:
fetch and FormData:
The ephemerality guarantee from the dashboard applies to scripted uploads too:
the audio is never stored, each request is discarded once every vendor
responds, and the run is recorded in your audit log as metadata (filename,
size, vendors compared) — never the audio. Re-run your reference clip set on a
schedule and the audit log stays a metadata trail you can correlate.
Clips are sent only to the catalogued STT vendors for transcription quality
comparison. They are never used to originate any call and never touch a PSTN
leg.
Reading the comparison
Read the three numbers together rather than picking the highest single score:- Transcript first. A vendor with a slightly lower confidence score but a word-perfect transcript on your real audio beats a confident-but-wrong one. Run the same clip you expect in production — proper nouns, numbers, and noisy callers are where vendors diverge.
- Confidence as a tiebreaker. Transcript-equal vendors: prefer the higher self-reported confidence.
- Latency for the live path. For real-time agents, the latency delta between vendors on the same clip is what a caller will feel. For batch or post-call transcription it barely matters.
vendors field narrowed to a shortlist once you are converging
— uploads are cheap, and per-vendor numbers from one run are not directly
comparable to another run because each run measures the vendor fresh.
Ephemeral by design
Nothing you upload is stored. The clip lives only for the duration of the comparison request and is discarded the moment every vendor responds server- side; on the client, the browser discards the clip when you replace it, discard it, or navigate away. The comparison is recorded in your account’s audit log as metadata (filename, size, vendors compared) — never the audio.Before you wire an agent
The vendor you pick here becomes the one your agent’s transcription depends on. Use the playground as the last gate before you set it:- Compare your real audio — greetings, IVR prompts, the phrases your callers say.
- Shortlist on transcript accuracy, then confidence, then latency.
- Pick the winner in the agent’s STT settings.
Available here is
wireable in the agent config today, and a Not yet available vendor there is
just as unavailable here.
The TTS analogue
The text-to-speech counterpart already exists: a listen-back preview on the same authoring surfaces that lets you hear a voice before assigning it (POST /api/v1/voice/tts-preview behind the dashboard’s voice pickers). Use both
together — audition the TTS voice, and benchmark the STT vendor — so the agent
you wire sounds right and hears right on the first call.
When a vendor row shows an error
One bad vendor never sinks the whole comparison — the failed vendor gets its own error row and the rest still come back. The error message in the row names the problem so you know what to check next.Why does a vendor row say it cannot transcribe?
The playground returns two codes on a failed vendor row. The code appears in theerror field of that vendor’s result entry.
STT_CREDENTIAL_REJECTED — the vendor rejected the credential you have on
file. The most common fix is to verify your BYO key under
Settings → Voice, or revoke it to fall back to the platform key. A revoked
key takes effect on the next comparison run.
STT_PROVIDER_UPSTREAM — the vendor’s own service returned an error. This
is a transient upstream issue, not something wrong with your account or
credentials. Retry the comparison in a few minutes.
Both codes surface on the row itself, not as a top-level request error — the
other vendors in the run are unaffected.
Which API does the Playground hit?
The dashboard Playground callsPOST /api/v1/voice/stt-preview, the same
endpoint you can script directly. See the endpoint contract
above for the full request shape, response fields, and error statuses.
Where are these terms defined?
The glossary defines STT (Speech-to-Text), TTS, and the other voice-pipeline terms referenced on this page.What if the error is not one of these codes?
Open the troubleshooting index and search for the error message shown on the row. Most STT errors fall into one of three buckets: a credential that needs attention, a transient upstream outage, or an input clip with no audible speech.See also
- IVR routing with NLU — speech intents run on the transcription layer you are choosing here.
- Voice queues — routing callers to agents once the audio pipeline is tuned.