Skip to main content
Browse documentation
Capabilities

Speech to Text

Transcribe spoken audio into text.

When speech-to-text is enabled, use the recognition models and options exposed by the configured runtime. Language coverage, timestamps, diarization, entity extraction, and streaming behavior are provider-dependent.

  • Confirm batch versus realtime support in the active runtime.
  • Validate timestamps, diarization, language support, latency, and limits against representative audio before production use.
MobBotLearning guide

I’ll help turn the documentation into the next action.

Use the current page as the source, then move directly into the product, API, workflow, or support path that matches what you are trying to accomplish.