Speech recognition, or speech-to-text (STT), transcribes speech into text. Models like Whisper have made it accurate, multilingual and runnable locally. It is the entry point of any voice agent: transcription quality conditions everything that follows.
In practice at Gensai
Gensai runs faster-whisper locally for its voice installations, such as the Oceans device, and AssemblyAI for Qualicontact's donation-call analysis.
Going beyond the definition?
From concept to project: Gensai builds custom, sovereign, GDPR-compliant AI solutions.
Related services: Chatbots & Voicebots