How do you choose a voice AI model? Latency, expressivity, sovereignty: Sébastien Lecocq's interview with Gradium

How do you choose a voice AI model? Latency, expressivity, sovereignty: Sébastien Lecocq's interview with Gradium

An artificial intelligence interviewing a human about voice AI: that is the format chosen by Gradium for its series of conversations with the teams using its models. On 5 October 2026, Sébastien Lecocq, co-founder of Gensai, answered "Marlo", Gradium's synthetic voice. Three minutes, in English.

Those three minutes say a lot about the state of voice AI in 2026, and above all about the criteria that matter once it goes into production: response speed, the expressivity of the voice, and control over where the models run. A look back at the interview, and at what it offers anyone evaluating a voice provider.

The interview published by Gradium on 5 October 2026 (in English, 3 minutes).

What happens when the interviewer is an AI?

Marlo asks her questions, rephrases, thanks, follows up. The exchange flows, and that is precisely the point: for a provider of synthetic voices, the fluidity of the conversation is the product itself. A two-second hesitation before every answer, a flat intonation, and the demonstration collapses. Gradium therefore takes the risk of proving it live, which says more than any brochure.

For Gensai, it is a format with a future. The voice AI solutions designed by the studio rely on the same chain: speech recognition (STT, speech-to-text), reasoning by a language model connected to the client's data, then speech synthesis (TTS, text-to-speech).

Why has voice become central to Gensai's projects?

Asked "why is voice AI central to your expertise?", Sébastien answers with customer demand: "our customers require more and more voice interaction to interact with their app". A kiosk in a public venue, a visitor guide, a phone reception: in all these cases, screen and keyboard are a detour, speech is the direct route.

Two families of projects illustrate this shift. On one side, mediation and reception installations, where a character answers visitors out loud. On the other, phone agents that pick up, inform and take messages for practices and small organisations. Two very different contexts, one shared requirement: the voice must sound natural and arrive fast.

How do you make an AI speak without Wi-Fi? The aquarium case

The most telling example of the interview is an aquarium in Val d'Europe, for which Gensai designed a sea creature that talks with the public about the ocean. The constraint was radical: no Internet connection on site. Calling a cloud speech synthesis (TTS) API, however excellent, was out of the question.

The answer was Phonon, Gradium's on-device text-to-speech (TTS) model. Around one hundred million parameters, fully offline execution on a simple CPU, a voice that stays natural. The complete installation, speech recognition (STT), language model, knowledge base and synthesis (TTS), runs on a Mac Mini, with no cloud and no pay-per-use subscription. The story of the project is told in How to make the oceans speak with frugal AI and in the Immersive voice experience case study.

Going on-device goes beyond the technical constraint. It settles in one move the confidentiality of exchanges, robustness in a busy venue and energy frugality. For museums and events, it is often the right architecture, not a fallback.

What makes a synthetic voice sound natural? Latency and expressivity

In the interview, Sébastien ties the two notions together: "what's actually making it expressive and natural is the latency". The delay between the end of a question and the start of the answer, which adds up transcription time (STT), the model's reasoning and the first sound produced by the TTS, is the first signal the user perceives. Too long, and they repeat their question or give up. Short enough, and the dialogue settles in, the voice's hesitations become breaths rather than bugs.

Expressivity comes next: intonation, rhythm, the ability to read a phone number, an acronym or an e-mail address correctly without stumbling. It is on these hard cases that models stand out in production, far more than on a demo sentence.

Gradium announces first audio in under 50 milliseconds for its latest TTS model, 47 ms at the median according to its own measurements, and first audio in about 200 ms on its real-time API. Those are the provider's figures. On Gensai's side, the rule is simpler: latency is measured on the actual installation, in the client's conditions, and tuned with the agent's speaking pace.

Why a French partner?

The third criterion Sébastien cites is sovereignty. Gradium was founded in September 2025 in Paris by the researchers behind Kyutai, the open research lab known for its real-time speech systems, including Moshi. Its founders, Neil Zeghidour, Laurent Mazaré, Olivier Teboul and Alexandre Défossez, authored a good part of the work that shaped modern generative audio, from neural audio codecs to audio language models.

Seven months after launch, in July 2026, the company extended its seed funding to 100 million dollars, with NVIDIA among its new investors, and opened an office in the San Francisco Bay Area. It offers its models as an API, on dedicated instances, self-hosted or on premises, with a zero data retention policy. For a studio that promises its clients responsible, sovereign and frugal AI, being able to choose where the voices run, down to fully on site, is a selection criterion, not a marketing line.

What stands out from the interview is the nature of the relationship. Gensai integrated Phonon while the model was still in beta, in a real and demanding installation, and discussed the hard cases with Gradium's team. Sébastien puts it his own way: the quality of the exchanges with the team and the support, and having "a company that's like that in Paris", are part of what makes Gradium "cool".

For Gensai, the partnership brings a reference voice layer, French, available from the cloud to the device, and direct access to the people who design the models. For Gradium, field feedback on uses where voice is not a gimmick but the heart of the experience: a kiosk talking to families in an aquarium, an agent answering a practice's phone. The interview published by Gradium is the public trace of it.

How do you evaluate a voice AI provider?

To the last question, "what would you tell someone evaluating a voice AI model?", Sébastien answers first: try it. His list, completed by the practice of Gensai's projects, fits in five points.

  • Try it on your own cases, not on the provider's demo: your business texts, proper names, numbers, acronyms.
  • Measure latency end to end, STT, model and TTS included, in real conditions, on the phone or on the kiosk, with the venue's network, not only on a developer's machine.
  • Listen to expressivity on real dialogues: interruptions, resumptions, open questions, language switching.
  • Check where the models run: hosting in France or in Europe, dedicated instance, self-hosting, and the existence of an on-device version when the venue requires it.
  • Test the support before signing: the quality of the team's answers on a hard case says a lot about what follows.

The voice layer is an architecture choice

It is tempting to treat speech synthesis as an interchangeable commodity, one line in a quote. The interview reminds us of the opposite: voice is what the user perceives first, and its choice commits latency, confidentiality, cost per use and the ability to work without a network. A voice model is therefore chosen the way one chooses a database or a hosting provider, with criteria, tests and a lasting relationship with the people who build it.

Going further: the Voice AI page details the voice agents, voicebots, callbots and avatars designed by Gensai, and the article AI voice agent: how it works describes the technical building blocks involved.

Frequently asked questions

What is an on-device voice model?

A speech synthesis (TTS) or recognition (STT) model that runs directly on the device, computer, kiosk, tablet or phone, without calling a remote server. It works without an Internet connection, lets no data leave the device and costs nothing per use. Phonon, Gradium's on-device model, fits in about one hundred million parameters and runs on a simple CPU.

What do STT and TTS mean?

STT (speech-to-text) is speech recognition: speech is transcribed into text so that the language model can process it. TTS (text-to-speech) is speech synthesis: the written answer is turned into a voice. A voice agent chains STT, reasoning and TTS, and the perceived latency is the sum of the three. Gradium provides both building blocks, as a real-time API or as an on-device version for TTS.

Why does latency matter so much in a voice agent?

Because the silence between the question and the answer is the first signal the user perceives. The shorter it is, the more natural the dialogue feels, and the more the voice's expressivity becomes audible. When it stretches, the user repeats, talks over the agent or hangs up. Latency is measured on the actual installation, in the venue's conditions, and tuned with the agent's speaking pace.

Who is Gradium?

A Paris company founded in September 2025 by the researchers behind the Kyutai lab. It develops real-time voice models (synthesis, recognition, live translation, voice cloning and design, on-device model) offered as an API, on dedicated instances, self-hosted or on premises. In July 2026 it extended its seed funding to 100 million dollars, with NVIDIA among its investors.

Can a voice AI be entirely French?

Yes, across the whole chain: speech recognition and synthesis with models designed in France, an open language model, hosting with a French provider or local execution on the device. It is the architecture Gensai chooses when confidentiality or the venue requires it, and a selection criterion from the moment providers are compared.

An AI project in mind?

Chatbot, AI agent, answer engine or a first quick win: Gensai reviews every need and suggests a suitable approach.

Related services: Voice AI

55 boulevard de Strasbourg, 75010 Paris
Information used only to answer the request.