Urgent.News

What's breaking now, across thousands of outlets.

AI

Google launches Gemini 3.5 Transcribe, which powers Gboard Rambler & is coming to Chrome

Google today introduced Gemini 3.5 Transcribe as its “most precise speech-to-text model yet” that is already powering several first-party products.

Google launches Gemini 3.5 Transcribe, which powers Gboard Rambler & is coming to Chrome

Google has unveiled Gemini 3.5 Transcribe, heralded as its most precise speech-to-text model to date. This cutting-edge technology is already integrated into a range of first-party products, including Spark, Gmail, and beyond. Unlike traditional speech recognition models that grapple with background noise, intricate terminology, and disfluency refinement, Gemini 3.5 Transcribe transforms raw audio into precise, polished, and formatted text.

This advanced model is engineered to capture the user's unique speaking style, enabling a more accurate comprehension of their intentions and recognition of custom vocabulary. Gemini 3.5 Transcribe can adeptly manage self-corrections like "let's meet Tuesday—no, Wednesday" and eliminate filler words such as "ums" and "ahs" from the final output. Moreover, it possesses the ability to automatically format text and facilitate real-time voice editing, enhancing the overall user experience.

Performance-wise, Gemini 3.5 Transcribe delivers a significant leap in capabilities compared to its predecessor, Chirp 3. It boasts superior word error rates and notably improved latency. According to Artificial Analysis, the time to finalize transcription sees a 70% enhancement. Additionally, the model demonstrates exceptional multilingual performance across various languages and locales, surpassing Chirp 3 on the FLEURS benchmark. It achieves a 5.50% word error rate in streaming mode, and a 5.04% WER in non-streaming use-cases.

The ultimate objective of Gemini 3.5 Transcribe is to empower users to execute tasks through voice commands. Function calling facilitates Gemini 3.5 Transcribe to delegate complex tasks, like image generation and file analysis, to other Gemini models. This functionality is exemplified by the "Speak to Window" feature in the Gemini app for macOS.

In addition to the Gemini macOS app and Gboard Rambler on Android, the model is set to launch in Google Antigravity's prompt box microphone. Here, it leverages screen context and chat history, with user permission, to ensure unparalleled transcription accuracy across file names, agent thoughts, and active documents. Gemini 3.5 Transcribe is soon to be available in the Chrome browser, allowing users to "talk to type" in any web field, making the process of dictating replies, drafting posts, or prompting Gemini in Chrome more seamless and natural.

Written by urgent.news from 9to5Google's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at 9to5google.com →

More in AI

Beyond the Hype: 4 Agentic Design Patterns Every Dev and PM Needs to Know

The current AI landscape is thick with "smoke." Between infinite buzzwords and thousands of AI posts and infographics, it is becoming increasingly difficult to discern what is actually a new…

  • The Pipeline treats tasks as specialized nodes with sequential inputs
  • The Router introduces branched logic for query analysis and delegation
  • The Evaluator-Optimizer uses adversarial loop for quality assurance and refinement

More from Wednesday 26 August →