Gemini can now clone your voice and perform scripts like an actor
The upgrade adds 2,000+ voices, two-speaker scenes, and finer delivery controls.
Google has introduced Gemini 3.8 Flash TTS and Flash-Lite TTS, cutting-edge text-to-speech models designed to produce more engaging AI-generated audio. These new models allow users to craft custom voices, recreate authorized voices from a 30-second sample, and fine-tune aspects like accents, pacing, emotion, and dialogue between two speakers.
Gemini 3.8 Flash TTS and Flash-Lite TTS are now accessible on Gemini Notebook and Google Vids, and developers can integrate them through Google AI Studio and the Gemini API. Users can feed the models a script, and the AI will deliver the audio, but achieving the desired tone—accurate accent, suitable pacing, appropriate emotion, and precise pauses—remains a challenging task.
To tackle this issue, Google has invested time in expanding Gemini 3.8 from the original Gemini 3.8 Flash into new Live and Live Extended Thinking models. The latest rollout now enables Gemini 3.8 Flash TTS and Flash-Lite TTS to generate original voices based on natural-language descriptions, supporting over 100 languages and dialects. Over 2,000 production-ready voices are available, and the technology can replicate a consistent voice from a 30-second sample, provided the appropriate usage permissions are in place.
Written by urgent.news from Android Authority's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.