Intelligent transcription with Gemini 3.5 Transcribe
Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.
Gemini Audio has unveiled Gemini 3.5 Transcribe, their latest speech-to-text model engineered for precise voice interactions. This model is capable of converting raw audio into accurate, polished, and formatted text, even in challenging conditions like background noise, technical jargon, and disfluencies. Users of the Gemini app and Android devices have already started reaping the benefits of this transcription model through new voice functionalities such as Rambler on Android and the Gemini app on macOS.
Developers can now integrate Gemini 3.5 Transcribe into their applications using the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. This new model can be effortlessly incorporated into existing workflows, whether the goal is to create voice agents, real-time captioning tools, or post-call analytics pipelines.
Gemini 3.5 Transcribe offers two versions of the API: one designed to capture users' natural speaking style and better understand their intent, including custom vocabulary for task execution, and another aimed at improving word error rates and latency. Performance improvements are evident, with a 70% reduction in time to final transcription and a 5.04% word error rate (WER) in non-streaming use-cases, according to tests conducted by Artificial Analysis.
The model's multilingual capabilities also surpass those of its predecessor, Chirp 3, delivering precise performance across various languages and locales. Furthermore, beyond traditional speech-to-text functionality, 3.5 Transcribe introduces context-aware understanding, making interactions across Google platforms feel more natural and intuitive. Features like Gboard, Antigravity, the Gemini app, and Chrome now leverage the model to capture nuances, intent, and inline edits with ease.
Developers can further streamline their processes by utilizing the Gemini Live API in conjunction with third-party platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms handle complex real-time media streaming infrastructure, enabling developers to concentrate on crafting user experiences. Companies such as Vivo, Intellitek Health, and Lingopal have praised 3.5 Transcribe for its impressive latency, accuracy, and extensive language support.
Written by urgent.news from Google DeepMind's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.