← AI Glossary

Speech synthesis (TTS)

Language, voice & vision

Speech synthesis, or text-to-speech (TTS), turns text into voice. Recent models produce natural, expressive, cloneable voices in many languages. In a voice agent, it provides the timbre and the embodiment; choosing the voice is as much a design decision as a technical one.

In practice at Gensai

For the Oceans, Gensai gave the sea a deep, mysterious voice through Gradium's speech synthesis.