# Speech to Text



> Transcribe spoken audio into text.



When speech-to-text is enabled, use the recognition models and options exposed by the configured runtime. Language coverage, timestamps, diarization, entity extraction, and streaming behavior are provider-dependent.

- Confirm batch versus realtime support in the active runtime.
- Validate timestamps, diarization, language support, latency, and limits against representative audio before production use.



---

Powered by MobDial Agents (MobAgents).

