Google launches two benchmark-topping speech generation models
Google LLC today made two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, available through its cloud platform. The algorithms have highly similar application programming interfaces, which makes using them side-by-side relatively simple for developers. Flash-Lite TTS is optimized for cost-cost efficiency and inference speed. Flash TTS offers better audio quality […]…
Google has unveiled two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which have outperformed competitors in benchmark tests. Both models are accessible through Google's cloud platform and share similar application programming interfaces. Flash-Lite TTS prioritizes cost-efficiency and speed, while Flash TTS provides superior audio quality at a higher price.
The models can generate speech in 130 languages on launch, compared to 101 for Flash-Lite TTS, and offer over 2,000 prepackaged voices. Custom voices can be created using natural language prompts, and developers can customize vocal timbre, accent, and pacing. Additionally, Google allows developers to generate synthetic voices based on a 30-second audio sample, requiring consent from the speaker.
Future customization options may enable creating new voices by modifying prepackaged ones. The models also feature customizable delivery, allowing developers to add oratory cues to each line of the script, resulting in non-lexical vocalizations and pacing shifts. Google employs a technology called SynthID to embed an audio watermark in the generated speech, which is undetectable to humans but can be identified by AI detection tools.
To further ensure authenticity, Google attaches a C2PA record to every generated audio file, specifying generation time, modifications, and related details. In benchmark tests conducted by startup Hume AI Inc., Flash TTS and Flash-Lite TTS secured the top two spots, outperforming several competing models and several language-specific versions of Voice Arena.
These models enable creators, developers, and enterprises to produce more expressive audio experiences, enhancing user experiences in products like Gemini Notebook and Google Vids.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.