Intelligent transcription with Gemini 3.5 Transcribe
Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.
Gemini Audio has unveiled its latest innovation, Gemini 3.5 Transcribe, an advanced speech-to-text model designed for intelligent voice interactions. This new model excels in converting raw audio into high-quality, formatted text, outperforming traditional systems that often falter with background noise, technical jargon, and disfluencies.
Developers can now integrate this powerful transcription tool into their products using the Gemini API within Google AI Studio, or via the Gemini Enterprise Agent Platform. Gemini 3.5 Transcribe aims to seamlessly integrate with a variety of applications, from voice agents and real-time captioning tools to post-call analytics pipelines.
The model offers enhanced customization by learning users' natural speaking styles and learning custom vocabulary, making it ideal for executing tasks with voice commands. Its performance marks a significant leap forward from the previous Chirp 3 model, boasting improved word error rates and reduced latency. According to Artificial Analysis, time to final transcription is 70% faster with Gemini 3.5 Transcribe.
The model also demonstrates exceptional multilingual capabilities, surpassing previous models on the FLEURS benchmark across various languages and locales, achieving a 5.50% word error rate in streaming mode and 5.04% in non-streaming cases.
Beyond its transcription capabilities, 3.5 Transcribe makes working with Google more intuitive by understanding context directly within everyday interfaces like Gboard, Antigravity, the Gemini app, and Chrome. It captures nuances, intent, and inline edits with ease. Developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents leverage the Gemini Live API to build and deploy high-performance voice-driven interfaces without worrying about complex real-time media streaming infrastructure.
Companies like Vivo, Intellitek Health, and Lingopal have praised 3.5 Transcribe for its impressive latency, accuracy, and extensive language support.
Written by urgent.news from Google Blog's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- Intelligent transcription with Gemini 3.5 Transcribe deepmind.google
- Here’s how to use intelligent dictation in Gemini for macOS. blog.google