Gemini 3.8 text-to-speech says hello
Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS are our most expressive audio models yet.
Google has unveiled two new text-to-speech models as part of their Gemini Audio family, designed to create realistic and customizable voices for a wide range of applications. These tools allow users to craft unique voices, adjusting aspects like accent and emotional tone, making them ideal for audiobooks, games, and podcasts that sound like real people speaking. To ensure responsible use, Google has also incorporated safety features into these new models.
The company has introduced two advanced text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which offer creators, developers, and businesses an enhanced toolkit for generating high-quality audio experiences. These models enable users to go beyond the preset options, creating dynamic and expressive audio content. By incorporating them into products like Gemini Notebook and Google Vids, Google aims to improve user experiences.
These models represent an expansion of Google's Gemini Audio family, which has seen significant growth since the launch of features such as Live Translate, Live Transcribe, Live, and Live Extended Thinking. The new models provide users with the capability to generate voices from a vast library of over 100 languages, allowing for the creation of entirely original character voices or consistent brand ambassadors.
To ensure responsible usage, Google has implemented stringent safeguards, including consent verification for voice replication, ensuring that a verbal consent recording from the voice owner matches the reference speaker before a voice can be created. Additionally, every audio clip generated by the Gemini Audio models is watermarked with SynthID, an imperceptible marker woven into the audio output, helping to prevent misinformation.
Google is offering developers early access to these new speech generation capabilities through Google AI Studio, a voice design workspace that allows users to create entirely new vocal identities or replicate their own voice. The models can be integrated into developer platforms such as Agora, LiveKit, Pipecat, and Vercel, enabling the creation of high-performance speech generation experiences.
Some of the companies integrating Google's latest TTS models include Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang, who are leveraging these tools to accelerate global dubbing, localize media with regional accents, and power conversational voice agents at scale. However, it's worth noting that voice replication through AI Studio is currently unavailable in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, and India.
Written by urgent.news from Google Blog's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
Also reported by 2 other outlets
- Gemini 3.8 text-to-speech says hello deepmind.google
- Gemini 3.8 text-to-speech says hello blog.google