Speech recognition, or speech-to-text (STT), transcribes speech into text. Models like Whisper have made it accurate, multilingual and runnable locally. It is the entry point of any voice agent: transcription quality conditions everything that follows.
In practice at Gensai
Gensai runs faster-whisper locally for its voice installations, including the Oceans device and donation-call analysis.