{
  "id": 3560377,
  "title": "Intelligent transcription with Gemini 3.5 Transcribe",
  "url": "https://urgent.news/2026/08/26/intelligent-transcription-with-gemini-3-5-transcribe-3560377",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-26T17:01:00.000Z",
  "source": {
    "name": "Google DeepMind",
    "slug": "google-deepmind",
    "url": "https://deepmind.google/blog/intelligent-transcription-with-gemini-3-5-transcribe/"
  },
  "original_language": "en",
  "account": "Gemini Audio has unveiled Gemini 3.5 Transcribe, their latest speech-to-text model engineered for precise voice interactions. This model is capable of converting raw audio into accurate, polished, and formatted text, even in challenging conditions like background noise, technical jargon, and disfluencies. Users of the Gemini app and Android devices have already started reaping the benefits of this transcription model through new voice functionalities such as Rambler on Android and the Gemini app on macOS.\n\nDevelopers can now integrate Gemini 3.5 Transcribe into their applications using the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. This new model can be effortlessly incorporated into existing workflows, whether the goal is to create voice agents, real-time captioning tools, or post-call analytics pipelines.\n\nGemini 3.5 Transcribe offers two versions of the API: one designed to capture users' natural speaking style and better understand their intent, including custom vocabulary for task execution, and another aimed at improving word error rates and latency. Performance improvements are evident, with a 70% reduction in time to final transcription and a 5.04% word error rate (WER) in non-streaming use-cases, according to tests conducted by Artificial Analysis.\n\nThe model's multilingual capabilities also surpass those of its predecessor, Chirp 3, delivering precise performance across various languages and locales. Furthermore, beyond traditional speech-to-text functionality, 3.5 Transcribe introduces context-aware understanding, making interactions across Google platforms feel more natural and intuitive. Features like Gboard, Antigravity, the Gemini app, and Chrome now leverage the model to capture nuances, intent, and inline edits with ease.\n\nDevelopers can further streamline their processes by utilizing the Gemini Live API in conjunction with third-party platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms handle complex real-time media streaming infrastructure, enabling developers to concentrate on crafting user experiences. Companies such as Vivo, Intellitek Health, and Lingopal have praised 3.5 Transcribe for its impressive latency, accuracy, and extensive language support.",
  "summary": "Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "Google Blog",
        "title": "Here’s how to use intelligent dictation in Gemini for macOS.",
        "url": "https://urgent.news/2026/08/25/heres-how-to-use-intelligent-dictation-in-gemini-for-macos",
        "published": "2026-08-25T16:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}