Urgent.News

What's breaking now, across thousands of outlets.

AI

Gemini Live audio

Tool: Gemini Live audio Google released Gemini 3.8 Live and 3.8 Live Extended Thinking today - two new speech-to-speech models that are a similar shape to OpenAI's GPT-Live family. I pointed GPT-6 Astra Extra High at the documentation and had it build me this web UI for trying out the new models. You can select a model and voice preset, enter an optional system prompt and then start a voice…

Google has unveiled Gemini 3.8 Live and 3.8 Live Extended Thinking, two new speech-to-speech models inspired by OpenAI's GPT-Live series. To test these models, a custom web user interface was developed, allowing users to choose a model and voice preset, input a system prompt if desired, and engage in a browser-based voice conversation.

The interface connects directly to Google's wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent WebSocket endpoint, utilizing the Web Audio API for audio capture and playback. A comprehensive tutorial on utilizing this WebSocket API was released by Simon Willison on September 15th, 2026.

Written by urgent.news from Simon Willison's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at simonwillison.net →

More in AI

More from Tuesday 15 September →