Urgent.News

What's breaking now, across thousands of outlets.

AI

Gemini-3.5-Transcribe

Gemini Audio announced the launch of Gemini 3.5 Transcribe, their most accurate speech-to-text model yet. This advanced model is specifically designed to handle voice interactions with exceptional precision, even in challenging environments with background noise, complex jargon, and disfluencies.

Gemini 3.5 Transcribe outperforms traditional speech recognition models by converting raw audio into polished, formatted text directly. This powerful tool is already enhancing user experiences across various Google products, such as the Gemini app for macOS and the Gemini app for Android, thanks to new voice capabilities like Rambler.

Developers can now integrate this cutting-edge technology into their own products using the Gemini API, available in the Google AI Studio and the Gemini Enterprise Agent Platform. The model seamlessly integrates into developer workflows, whether for creating voice agents, real-time captioning tools, or post-call analytics pipelines.

Gemini 3.5 Transcribe offers two separate APIs, each with unique capabilities. The first is designed to capture individual speaking styles, understand custom vocabulary, and execute tasks with voice commands. The second API demonstrates significant improvements over the previous Chirp 3 model, with a major advancement in performance, including improved word error rates and reduced latency. According to Artificial Analysis, time to final transcription has improved by 70%.

Moreover, Gemini 3.5 Transcribe delivers exceptional multilingual performance, outperforming Chirp 3 on the FLEURS benchmark across various languages and locales. It achieves a 5.50% word error rate in streaming mode and a 5.04% word error rate in non-streaming use-cases.

Beyond the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, 3.5 Transcribe goes a step further by enhancing everyday interactions on Google platforms like Gboard, Antigravity, the Gemini app, and Chrome. By leveraging the Gemini Live API, developers can now build high-performance voice-driven interfaces using platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents.

These platforms handle complex real-time media streaming infrastructure, allowing developers to focus on crafting an intuitive user experience.

Companies like vivo, Intellitek Health, and Lingopal have already praised Gemini 3.5 Transcribe for its impressive latency, accuracy, and extensive language support.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at blog.google →

More in AI

More from Thursday 27 August →