{
  "id": 7614144,
  "title": "Google Introduces Gemini 3.8 Live Audio Models for Real-Time Voice AI Workflows",
  "url": "https://urgent.news/2026/09/15/google-introduces-gemini-3-8-live-audio-models-for-real-time-voice-ai",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-15T19:15:30.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/alifar/google-introduces-gemini-38-live-audio-models-for-real-time-voice-ai-workflows-5cfd"
  },
  "original_language": "en",
  "account": "Google DeepMind has unveiled two new audio-focused models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed for more natural, near real-time conversations with AI. These models expand beyond the existing transcription, translation, and text-to-speech capabilities offered by the Gemini Audio family.\n\nGemini 3.8 Live is geared towards high-volume, cost-effective voice interactions, offering conversational capabilities optimized for real-time use. This model is particularly well-suited for applications requiring quick and responsive voice interactions, such as customer-facing services or internal workflows. On the other hand, Gemini 3.8 Live Extended Thinking is designed for more complex reasoning tasks, enabling the orchestration of multiple agents to handle background tasks while maintaining a natural conversation flow.\n\nFor businesses, the introduction of these models opens up new possibilities for voice AI, allowing for more practical and intuitive interfaces for various tasks. Rather than relying solely on typing instructions, voice systems can now support conversational task handling, customer interactions, and guided workflows. However, it's important to note that Google has not yet published explicit public pricing for either model, leaving the operational comparison to be determined by individual businesses.\n\nGemini 3.8 Live and Extended Thinking are part of the broader Gemini Audio portfolio, which also includes models like Gemini 3.5 Transcribe, Gemini 3.5 Live Translate, and Gemini 3.1 Flash TTS. These models serve distinct audio tasks, such as converting speech to text, translating live audio, and generating speech. The new live models, however, focus on the conversational layer itself, providing real-time dialogue and reasoning capabilities.\n\nGoogle offers multiple access routes to these new models, including Google AI Studio, the Gemini Live API, the Gemini API, the Gemini app, and related Google services. This indicates that the Gemini Audio capabilities are intended to cater to a range of experimentation, development, and end-user experiences, rather than being limited to a single product.\n\nWhile the official documentation does not provide explicit public pricing or detailed operational information, such as token costs and latency figures, organizations considering voice AI deployment should consider several factors before making a decision. These considerations include the specific workflow requirements, the access path that best fits the intended product or internal process, the integration with existing business systems, and the need for testing to ensure the quality of responses for real conversations and tasks.\n\nUltimately, voice AI has the potential to streamline tasks that begin with a spoken request, but a successful deployment requires a clear connection between the conversation and the subsequent action. This may involve routing information to existing applications, triggering approved workflows, or involving the user when a decision requires review. The multi-agent capability of Gemini 3.8 Live Extended Thinking is particularly noteworthy, but its practical value will depend on how effectively it can be integrated with the specific tools, data, and steps involved in the underlying work.",
  "summary": "Google DeepMind has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking , two audio-focused models designed for more natural, near real-time conversations with AI. The models expand the Gemini Audio family beyond transcription, translation, and text-to-speech capabilities, with one aimed at high-volume voice interactions and the other built for more demanding reasoning while a…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "Google Blog",
        "title": "Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe",
        "url": "https://urgent.news/2026/09/15/build-real-time-voice-applications-with-gemini-3-8-live-and-3-5",
        "published": "2026-09-15T17:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}