Capabilities
Speech to Text
Transcribe spoken audio into text.
When speech-to-text is enabled, use the recognition models and options exposed by the configured runtime. Language coverage, timestamps, diarization, entity extraction, and streaming behavior are provider-dependent.
- Confirm batch versus realtime support in the active runtime.
- Validate timestamps, diarization, language support, latency, and limits against representative audio before production use.
Machine-readable source: /docs/overview/capabilities/speech-to-text.md