Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI makes you call sales for a custom voice. Google just made it self-serve.

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS today through the Gemini API and Google AI Studio. The post OpenAI makes you call sales for a custom voice. Google just made it self-serve. appeared first on The New Stack .

OpenAI makes you call sales for a custom voice. Google just made it self-serve.

Google has released two new text-to-speech options through its Gemini API and Google AI Studio: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Unlike previous text-to-speech APIs, these new offerings allow developers to create their own custom voices. Users can describe the voice they desire or provide a short recording of an existing voice, which Google then uses to generate a unique voice ID.

This ID can be saved and reused throughout an application, eliminating the need for users to provide recordings or descriptions with every new request.

To create a voice, developers must submit two audio recordings from the same speaker, each lasting between 10 and 30 seconds. Additionally, a consent recording must be provided, where the speaker confirms that the voice belongs to them and agrees to allow Google to create a synthetic version. Google verifies that the person giving consent and the reference clip are indeed the same before proceeding with voice creation.

Once approved, Google returns a voice_… ID and stores it in the developer's project for a year. Developers can manage their voices through the API, retrieving, listing, or deleting them as needed. A project can hold up to 200 voices in total. Developers can also choose to use voice replication without storing the profile in the project, which returns an encrypted voicekey_… instead. This option is suitable for short-lived jobs, as the key expires after seven days.

Google has taken steps to ensure transparency and traceability with its new voice replication feature. Generated audio from Gemini is marked with SynthID, and replicated voices carry C2PA content credentials, allowing users to trace the origin of the audio. However, Google does not offer voice replication through AI Studio in certain regions, including Illinois, Texas, the European Economic Area, the U.K., Switzerland, and India.

Developers can create custom voices using natural language descriptions, specifying role, accent, and character. The system supports over 100 languages and dialects for Flash TTS and 101 for Flash-Lite. Google also offers a library of more than 2,000 prebuilt voices, with 30 studio voices and hundreds more in an extended library that can be filtered by language, accent, pitch, and use case. A remixing feature for adjusting timbre, pitch, pace, and accent is also in the works.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 3 other outlets

Read the original at thenewstack.io →

More in AI

More from Thursday 24 September →