Urgent.News

What's breaking now, across thousands of outlets.

AI

Google debuts Gemini 3.5 Transcribe, a speech-to-text model that powers Gboard Rambler and is coming to Chrome, in public preview for developers and enterprises (Abner Li/9to5Google)

Google has unveiled Gemini 3.5 Transcribe, a cutting-edge speech-to-text model that demonstrates remarkable precision. This model is already integrated into several first-party products, including Google's Gemini Live, which is set to receive productivity enhancements through Spark, Gmail, and other integrations.

What sets Gemini 3.5 Transcribe apart from traditional speech recognition models is its ability to effectively manage background noise, complex jargon, and disfluencies. It seamlessly converts raw audio into accurate, polished, and formatted text, while also capturing users' natural speaking styles to better understand their intentions and recognize custom vocabulary.

The model adeptly handles self-corrections, such as "let's meet Tuesday—no, Wednesday," and eliminates filler words like "ums" and "ahs" from the final output. Furthermore, it auto-formats text and facilitates natural voice editing.

In terms of performance, Gemini 3.5 Transcribe represents a "major advancement" over its Chirp 3 transcription model from 2025. Artificial Analysis reports a 70% improvement in time to final transcription, and according to the FLEURS benchmark, the model exhibits superior multilingual performance across various languages and locales. It achieves a 5.50% Word Error Rate (WER) in streaming mode and a 5.04% WER in non-streaming use-cases, outperforming Chirp 3.

An additional feature of Gemini 3.5 Transcribe is its "function calling" capability, which enables it to delegate complex tasks, such as image generation and file analysis, to other Gemini models. This functionality is exemplified through the "Speak to Window" feature in the Gemini app for macOS. Beyond the Gemini macOS app and Gboard Rambler on Android, the model is also accessible in Google Antigravity's prompt box microphone.

In this application, it leverages screen context and chat history, with user permission, to ensure the utmost transcription accuracy across file names, agent thoughts, and active documents.

Google plans to extend Gemini 3.5 Transcribe to the Chrome browser, allowing users to "talk to type in any web field." This enhancement will make it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with voice commands. To learn more about this development, check out 9to5Google on YouTube.

Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at 9to5google.com →

More in AI

More from Wednesday 26 August →