Gemini 3.8 Live and 3.8 Live Extended Thinking
Google has introduced two new AI models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed to make voice interactions more natural and intelligent. These models excel in complex reasoning, real-time visual context, and background task execution without disrupting conversations. Users can now access these features immediately through the Gemini API, Google Workspace, and the Gemini app.
The new models are capable of handling interruptions, switching between languages, and explaining their thought process as they work. They provide advanced near real-time reasoning, enhancing voice agents and making conversations with AI more intuitive and intelligent. For developers and enterprises, these models offer the foundation for reliable, production-ready voice agents, making voice conversations with Gemini across various platforms smoother and more collaborative.
Gemini 3.8 Live Extended Thinking stands out with its enterprise-grade task completion and intelligence, ranking #1 overall in the Speech to Speech Quality Index (82.6) and leading in agentic task completion. It also demonstrates strong reasoning capabilities, achieving a high score of 97.7% on Big Bench Audio. Additionally, Gemini 3.8 Live has garnered a strong preference among users, securing second place in the Speech Agent Arena.
The models are not only highly effective but also cost-efficient, providing developers and enterprises with a scalable and capable solution.
Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses. It can automatically detect and transition between 97 supported languages mid-conversation. The models can execute tools and API calls in the background while maintaining the conversation, enabling natural acknowledgments of requests and keeping the dialogue flowing even when tasks are being processed in the background.
Gemini 3.8 Live Extended Thinking reasons and speaks simultaneously, providing deeper intelligence for complex workflows while ensuring an uninterrupted conversational flow. It uses natural verbal cues like "Let me check that…" to acknowledge prompts and offers live progress narration for multi-step background tasks.
The new models are integrated into Google Workspace and Search, offering more intuitive and collaborative experiences, particularly when dealing with complex tasks. Through the Gemini Live API, platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to create and deploy high-performance voice-driven interfaces effortlessly. These platforms manage complex real-time media streaming infrastructure, allowing developers to focus on the user experience.
To ensure transparency and prevent misinformation, all audio generated by Google's AI products is watermarked with SynthID. This watermark is barely perceptible, directly embedded in the audio output, allowing AI-generated content to be detectable and aiding in the prevention of misinformation. For further information on Google's approach to safety and responsibility, refer to the model card. Gemini 3.8 Live Extended Thinking is now available starting today.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
Also reported by 2 other outlets
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking deepmind.google
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking blog.google