Google's New Voice AI Tops One Ranking, Trails on Others
Gemini 3.8 Flash TTS topped a voice benchmark run by a company that also licenses to Google, but trails on matching a cloned voice to its speaker and on following volume instructions.
Estimated reading time: 4 minutes
TL;DR
- Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026, for generating speech from text.
- The models support custom voices and instructions for tone, pacing and delivery. Voice replication requires a consent recording from the voice owner.
- In Hume AI’s tests, Flash TTS ranked first for voice design, but its results varied across tasks, including weaker performance in matching a cloned voice to its original speaker.
- At paid Standard rates listed on September 28, audio generation costs approximately 1.4 US cents per minute for Flash TTS and 0.9 cents for Flash-Lite TTS through December 31. Text processing costs extra.
What happened
Google has released two models that turn written scripts into spoken audio, with instructions controlling how lines are delivered. Gemini 3.8 Flash TTS and Flash-Lite TTS began rolling out on September 23 through the Gemini API and Google AI Studio, according to Google’s announcement.
Google positions Flash TTS for narration, expressive performances and complex dialogue, and Flash-Lite TTS for higher-volume speech generation. Its developer documentation lists support for 130 languages in Flash TTS and 101 in Flash-Lite TTS. Both support preset voices, voices created from descriptions and voice replication.
Google says users can replicate a voice from a 30-second sample. This requires a verbal consent recording from the voice owner that matches the reference speaker. Replication through AI Studio is unavailable in Illinois, Texas, the European Economic Area, the UK, Switzerland and India.
The company also says both models can follow directions for pacing and emotion, generate two-speaker scenes and include sounds such as laughter and sighs. Generated audio carries an inaudible SynthID watermark; voice replication also uses C2PA metadata to record information about the content’s origin.
Google’s pricing table, checked on September 28, lists paid Standard audio output at $9 per million tokens for Flash TTS and $6 for Flash-Lite TTS through December 31, 2026. At the published rate of 25 audio tokens per second, that is approximately 1.4 and 0.9 US cents per minute respectively. Text input is billed separately, and the listed audio rates double from January 1, 2027.
What this means (and what it does not)
The models give developers controls for both the voice they generate and its delivery. These serve different purposes: creating a fictional narrator, for example, does not require reproducing an existing speaker accurately.
That distinction appears in Hume AI’s evaluation. Flash TTS ranked first overall for voice design, but 11th among 13 models for similarity to the original speaker in voice replication, scoring 3.53 out of five. In a separate test of following volume instructions, its pass rate was 56.9%, compared with 88.9% for Inworld Voice Design.
Hume tested preview versions. It says each audio clip received three human ratings, with model identities hidden and samples presented in random order. Hume discloses a non-exclusive licensing agreement with Google and says it does not share its reserved test materials with model development teams.
These results support a comparison of specific capabilities, rather than a claim that one model is best for every speech-generation task. Voice-design performance alone does not establish cloning accuracy or precise control over every aspect of delivery.
What we still do not know
The launch announcement and model documentation reviewed do not give a measured response time under specified conditions. That leaves developers without a clear basis in those sources for estimating conversational delays in their own applications.
The cited evaluation also does not establish whether its preview-version results carry over unchanged to the released models. Nor do the overall rankings establish equivalent performance in every supported language or dialect.
Sources & Bylines
Every source cited in this article, gathered in one place.
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/ — Leland Rechis, Alan Cowen
- https://www.hume.ai/blog/newly-released-google-s-gemini-3-8-flash-tts-tops-hume-s-real-world-voiceeq-leaderboard
- https://ai.google.dev/gemini-api/docs/pricing
- https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts
Editorial check, counted automatically
- 4 sources cited
- 12 inline-linked claims
- 0 unsourced claims found
- 0 banned words found
- 2 numbers without context
Also available in Portugues (BR)