An artificial intelligence interviewing a human about voice AI: that is the format chosen by Gradium for its series of conversations with the teams using its models. On 5 October 2026, Sébastien Lecocq, co-founder of Gensai, answered "Marlo", Gradium's synthetic voice. Three minutes, in English.
Those three minutes say a lot about the state of voice AI in 2026, and above all about the criteria that matter once it goes into production: response speed, the expressivity of the voice, and control over where the models run. A look back at the interview, and at what it offers anyone evaluating a voice provider.
What happens when the interviewer is an AI?
Marlo asks her questions, rephrases, thanks, follows up. The exchange flows, and that is precisely the point: for a provider of synthetic voices, the fluidity of the conversation is the product itself. A two-second hesitation before every answer, a flat intonation, and the demonstration collapses. Gradium therefore takes the risk of proving it live, which says more than any brochure.
For Gensai, it is a format with a future. The voice AI solutions designed by the studio rely on the same chain: speech recognition (STT, speech-to-text), reasoning by a language model connected to the client's data, then speech synthesis (TTS, text-to-speech).
Why has voice become central to Gensai's projects?
Asked "why is voice AI central to your expertise?", Sébastien answers with customer demand: "our customers require more and more voice interaction to interact with their app". A kiosk in a public venue, a visitor guide, a phone reception: in all these cases, screen and keyboard are a detour, speech is the direct route.
Two families of projects illustrate this shift. On one side, mediation and reception installations, where a character answers visitors out loud. On the other, phone agents that pick up, inform and take messages for practices and small organisations. Two very different contexts, one shared requirement: the voice must sound natural and arrive fast.
How do you make an AI speak without Wi-Fi? The aquarium case
The most telling example of the interview is an aquarium in Val d'Europe, for which Gensai designed a sea creature that talks with the public about the ocean. The constraint was radical: no Internet connection on site. Calling a cloud speech synthesis (TTS) API, however excellent, was out of the question.
The answer was Phonon, Gradium's on-device text-to-speech (TTS) model. Around one hundred million parameters, fully offline execution on a simple CPU, a voice that stays natural. The complete installation, speech recognition (STT), language model, knowledge base and synthesis (TTS), runs on a Mac Mini, with no cloud and no pay-per-use subscription. The story of the project is told in How to make the oceans speak with frugal AI and in the Immersive voice experience case study.
Going on-device goes beyond the technical constraint. It settles in one move the confidentiality of exchanges, robustness in a busy venue and energy frugality. For museums and events, it is often the right architecture, not a fallback.
What makes a synthetic voice sound natural? Latency and expressivity
In the interview, Sébastien ties the two notions together: "what's actually making it expressive and natural is the latency". The delay between the end of a question and the start of the answer, which adds up transcription time (STT), the model's reasoning and the first sound produced by the TTS, is the first signal the user perceives. Too long, and they repeat their question or give up. Short enough, and the dialogue settles in, the voice's hesitations become breaths rather than bugs.
Expressivity comes next: intonation, rhythm, the ability to read a phone number, an acronym or an e-mail address correctly without stumbling. It is on these hard cases that models stand out in production, far more than on a demo sentence.
Gradium announces first audio in under 50 milliseconds for its latest TTS model, 47 ms at the median according to its own measurements, and first audio in about 200 ms on its real-time API. Those are the provider's figures. On Gensai's side, the rule is simpler: latency is measured on the actual installation, in the client's conditions, and tuned with the agent's speaking pace.
Why a French partner?
The third criterion Sébastien cites is sovereignty. Gradium was founded in September 2025 in Paris by the researchers behind Kyutai, the open research lab known for its real-time speech systems, including Moshi. Its founders, Neil Zeghidour, Laurent Mazaré, Olivier Teboul and Alexandre Défossez, authored a good part of the work that shaped modern generative audio, from neural audio codecs to audio language models.
Seven months after launch, in July 2026, the company extended its seed funding to 100 million dollars, with NVIDIA among its new investors, and opened an office in the San Francisco Bay Area. It offers its models as an API, on dedicated instances, self-hosted or on premises, with a zero data retention policy. For a studio that promises its clients responsible, sovereign and frugal AI, being able to choose where the voices run, down to fully on site, is a selection criterion, not a marketing line.
What stands out from the interview is the nature of the relationship. Gensai integrated Phonon while the model was still in beta, in a real and demanding installation, and discussed the hard cases with Gradium's team. Sébastien puts it his own way: the quality of the exchanges with the team and the support, and having "a company that's like that in Paris", are part of what makes Gradium "cool".
For Gensai, the partnership brings a reference voice layer, French, available from the cloud to the device, and direct access to the people who design the models. For Gradium, field feedback on uses where voice is not a gimmick but the heart of the experience: a kiosk talking to families in an aquarium, an agent answering a practice's phone. The interview published by Gradium is the public trace of it.
How do you evaluate a voice AI provider?
To the last question, "what would you tell someone evaluating a voice AI model?", Sébastien answers first: try it. His list, completed by the practice of Gensai's projects, fits in five points.
- Try it on your own cases, not on the provider's demo: your business texts, proper names, numbers, acronyms.
- Measure latency end to end, STT, model and TTS included, in real conditions, on the phone or on the kiosk, with the venue's network, not only on a developer's machine.
- Listen to expressivity on real dialogues: interruptions, resumptions, open questions, language switching.
- Check where the models run: hosting in France or in Europe, dedicated instance, self-hosting, and the existence of an on-device version when the venue requires it.
- Test the support before signing: the quality of the team's answers on a hard case says a lot about what follows.
The voice layer is an architecture choice
It is tempting to treat speech synthesis as an interchangeable commodity, one line in a quote. The interview reminds us of the opposite: voice is what the user perceives first, and its choice commits latency, confidentiality, cost per use and the ability to work without a network. A voice model is therefore chosen the way one chooses a database or a hosting provider, with criteria, tests and a lasting relationship with the people who build it.
Going further: the Voice AI page details the voice agents, voicebots, callbots and avatars designed by Gensai, and the article AI voice agent: how it works describes the technical building blocks involved.
