Skip to main content

STT Playground

Picking a speech-to-text vendor for a voice agent is a transcription-quality decision, and it is hard to make blind. The STT Playground lets you try a short clip against every speech-to-text vendor Orbit recognises — from the dashboard, with no API integration — and read each vendor’s transcript, confidence, and latency side by side before you commit one to an agent. Open it under Voice → STT Playground in the dashboard. It is gated to the owner, admin, and developer roles because each comparison run invokes metered upstream transcription calls. If a BYO STT key is rejected while you configure voice inference, see Troubleshooting: voice inference credential rejected.

What it compares

Each run sends your clip to every catalogued STT vendor independently and returns one row per vendor: Vendors are compared in parallel. If one vendor fails — bad credentials, an upstream outage — it gets its own error row and the rest of the comparison still comes back. One bad vendor never sinks the run.

Honest availability

Today, Deepgram is the vendor with a provisioned batch transcription path. The other catalogued vendors — Soniox, Speechmatics, AssemblyAI, and xAI — are recognised by Orbit and appear in the comparison, but are not yet provisioned. They show as “Not yet available” rather than a fabricated result. That honesty is deliberate: the playground shows you the same availability the per-agent STT picker sees, so the comparison never promises a vendor you cannot actually wire. A vendor’s status flips to Available on its own once its upstream credentials are provisioned.

Record or upload

Two ways to get a clip in — both feed the same comparison:
  • Record from your microphone. Use the mic button on the page; speak the kind of prompt a caller would actually say (numbers, names, accents, background noise — whatever your agents will face).
  • Upload a clip. Accepted formats are MP3, WAV, OGG, AAC, and M4A, up to 25 MB. Keep clips short and representative: one prompt, not a whole call.
Once a clip is in, you get an inline audio player to check it, then Compare providers to run the comparison. Discard the clip to start over.

Upload contract

The page calls a single endpoint, and the same contract applies if you script your own comparison: POST /api/v1/voice/stt-preview
  • multipart/form-data with one file field holding the audio clip.
  • Optional vendors form field: a comma-separated list of vendor ids to restrict the run (for example deepgram). Omit it to compare every catalogued vendor.
  • Auth: an API key or session token with the voice:write scope, on an owner, admin, or developer role.
  • Response: { results: [{ vendor, label, provisioned, transcript?, confidence?, latency_ms?, error? }] } — one entry per vendor compared.
Errors are specific: 415 when the upload is not multipart or not an accepted audio type, 422 when no file (or an empty file) was provided, and 413 when the clip exceeds the 25 MB limit.

Script your own comparison

The playground page is the same POST /api/v1/voice/stt-preview endpoint you can call yourself, so a comparison run drops straight into a benchmark script or a CI check that re-tests your reference clips. Upload a clip with curl — one file form field, and an optional vendors field to narrow the run:
Omit --form vendors=... to compare every catalogued vendor. A 200 response returns one entry per vendor compared. An unprovisioned vendor comes back as its own provisioned: false entry rather than failing the run, and an upstream failure lands on that vendor’s error field alone — one bad vendor never sinks the rest:
The same multipart body is easy to build programmatically with Node’s built-in fetch and FormData:
Errors are specific — check the status before reading results: The ephemerality guarantee from the dashboard applies to scripted uploads too: the audio is never stored, each request is discarded once every vendor responds, and the run is recorded in your audit log as metadata (filename, size, vendors compared) — never the audio. Re-run your reference clip set on a schedule and the audit log stays a metadata trail you can correlate.
Clips are sent only to the catalogued STT vendors for transcription quality comparison. They are never used to originate any call and never touch a PSTN leg.

Reading the comparison

Read the three numbers together rather than picking the highest single score:
  1. Transcript first. A vendor with a slightly lower confidence score but a word-perfect transcript on your real audio beats a confident-but-wrong one. Run the same clip you expect in production — proper nouns, numbers, and noisy callers are where vendors diverge.
  2. Confidence as a tiebreaker. Transcript-equal vendors: prefer the higher self-reported confidence.
  3. Latency for the live path. For real-time agents, the latency delta between vendors on the same clip is what a caller will feel. For batch or post-call transcription it barely matters.
Re-run with the vendors field narrowed to a shortlist once you are converging — uploads are cheap, and per-vendor numbers from one run are not directly comparable to another run because each run measures the vendor fresh.

Ephemeral by design

Nothing you upload is stored. The clip lives only for the duration of the comparison request and is discarded the moment every vendor responds server- side; on the client, the browser discards the clip when you replace it, discard it, or navigate away. The comparison is recorded in your account’s audit log as metadata (filename, size, vendors compared) — never the audio.

Before you wire an agent

The vendor you pick here becomes the one your agent’s transcription depends on. Use the playground as the last gate before you set it:
  1. Compare your real audio — greetings, IVR prompts, the phrases your callers say.
  2. Shortlist on transcript accuracy, then confidence, then latency.
  3. Pick the winner in the agent’s STT settings.
The same catalogue powers both, so a vendor that is Available here is wireable in the agent config today, and a Not yet available vendor there is just as unavailable here.

The TTS analogue

The text-to-speech counterpart already exists: a listen-back preview on the same authoring surfaces that lets you hear a voice before assigning it (POST /api/v1/voice/tts-preview behind the dashboard’s voice pickers). Use both together — audition the TTS voice, and benchmark the STT vendor — so the agent you wire sounds right and hears right on the first call.

When a vendor row shows an error

One bad vendor never sinks the whole comparison — the failed vendor gets its own error row and the rest still come back. The error message in the row names the problem so you know what to check next.

Why does a vendor row say it cannot transcribe?

The playground returns two codes on a failed vendor row. The code appears in the error field of that vendor’s result entry. STT_CREDENTIAL_REJECTED — the vendor rejected the credential you have on file. The most common fix is to verify your BYO key under Settings → Voice, or revoke it to fall back to the platform key. A revoked key takes effect on the next comparison run. STT_PROVIDER_UPSTREAM — the vendor’s own service returned an error. This is a transient upstream issue, not something wrong with your account or credentials. Retry the comparison in a few minutes. Both codes surface on the row itself, not as a top-level request error — the other vendors in the run are unaffected.

Which API does the Playground hit?

The dashboard Playground calls POST /api/v1/voice/stt-preview, the same endpoint you can script directly. See the endpoint contract above for the full request shape, response fields, and error statuses.

Where are these terms defined?

The glossary defines STT (Speech-to-Text), TTS, and the other voice-pipeline terms referenced on this page.

What if the error is not one of these codes?

Open the troubleshooting index and search for the error message shown on the row. Most STT errors fall into one of three buckets: a credential that needs attention, a transient upstream outage, or an input clip with no audible speech.

See also

  • IVR routing with NLU — speech intents run on the transcription layer you are choosing here.
  • Voice queues — routing callers to agents once the audio pipeline is tuned.