{
  "id": 7592851,
  "title": "Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe",
  "url": "https://urgent.news/2026/09/15/build-real-time-voice-applications-with-gemini-3-8-live-and-3-5",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-15T17:00:00.000Z",
  "source": {
    "name": "Google Blog",
    "slug": "google-blog",
    "url": "https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/"
  },
  "original_language": "en",
  "account": "Google has launched new Gemini Live models, Gemini 3.8 Live and 3.8 Live Extended Thinking, in the Gemini API and Google AI Studio. These models are designed for building real-time voice applications, allowing developers to create voice agents that can reason and execute tasks while maintaining the flow of conversations.\n\nGemini 3.8 Live represents a significant leap forward, capable of performing tasks during dialogue. For complex requests, the Extended Thinking version offers deeper reasoning, ranking at the top of Artificial Analysis’ leaderboard. On the other hand, Gemini 3.5 Transcribe is a dedicated speech-to-text model providing highly precise transcription across 85+ languages.\n\nThe 3.5 Transcribe model achieved an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming, enabling features such as sub-second captioning, call center agents, and real-time audio analytics. It supports 85+ languages, ensuring a broad range of voice experiences.\n\nDevelopers can access these models through various partners like Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents, which handle media streaming infrastructure for real-world deployment. Pricing for these models is competitive, with a cost of $0.005 per minute for audio input and $0.018 per minute for audio output.\n\nGoogle encourages developers to explore the new models through ai.studio/live, clone example apps from GitHub, or integrate the live API skill into their own agents. The company also invites developers to create audio experiences using their speech and music generation models, all available in the Gemini API.",
  "summary": "A look at how developers can build with our latest audio models, Gemini 3.8 Live, 3.8 Live Extended Thinking, and 3.5 Transcribe.",
  "key_points": [
    "Google launches Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps",
    "Gemini 3.8 Live excels in task performance during conversation",
    "3.5 Transcribe offers 85+ language support and sub-second captioning"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 3,
    "also_reported_by": [
      {
        "outlet": "Techmeme",
        "title": "Google launches Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its \"most advanced live dialogue models yet\", to more effectively enable voice agents (Google)",
        "url": "https://urgent.news/2026/09/15/google-launches-gemini-3-8-live-and-gemini-3-8-live-extended-thinking",
        "published": "2026-09-15T18:20:02.000Z"
      },
      {
        "outlet": "Dev.to",
        "title": "Google Introduces Gemini 3.8 Live Audio Models for Real-Time Voice AI Workflows",
        "url": "https://urgent.news/2026/09/15/google-introduces-gemini-3-8-live-audio-models-for-real-time-voice-ai",
        "published": "2026-09-15T19:15:30.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}