Urgent.News

What's breaking now, across thousands of outlets.

AI

Microsoft targets ultra-realistic voice agents with its first streaming transcription model

Microsoft Corp. today expanded its MAI artificial intelligence model family with its first streaming transcription model, debuting alongside two others focused on text-to-speech. They’re designed for developers who want to build voice agents that can listen to people’s voices and reply instantly, similar to how humans talk to one another. The most consequential of the […] The post Microsoft…

Microsoft targets ultra-realistic voice agents with its first streaming transcription model

Microsoft has launched its first streaming transcription model, MAI-Transcribe-2-Streaming, as part of its expanding MAI artificial intelligence model family. This model is designed to enable developers to create voice agents that can listen to human speech and respond instantly, similar to natural conversation. The model accepts human speech via a WebSocket, continuously updating the transcript as the person keeps talking, and confirms when the transcript is final.

This allows applications to display live captions or process a user's request before they have finished speaking.

MAI-Transcribe-2-Streaming is priced at $0.54 per audio hour, supports over 60 languages, and can automatically detect the language being spoken. It delivers first transcript hypotheses within an average of 320 milliseconds, though response speed may vary depending on network connection and the AI system generating the response.

The company also introduced two additional models focused on text-to-speech generation: MAI-Voice-2.1, which produces more expressive and high-fidelity outputs, and MAI-Voice-2.1-Flash, which offers faster response times and lower costs. Both models support 23 languages.

These new models demonstrate Microsoft's accelerated transition away from model providers like OpenAI and Anthropic, despite being a major investor in both. Microsoft AI Chief Executive Mustafa Suleyman has been concerned about the costs associated with these powerful frontier models and has instructed AI researchers to focus on developing the MAI model family to reduce costs and power Copilot agents in platforms like Excel and Outlook.

With this launch, Microsoft now has the necessary components to build powerful voice AI agents that can converse with users in a natural, humanlike manner. Voice agents require the ability to recognize and understand speech, decide on actions based on what was said, and generate audible responses. MAI-Transcribe-2-Streaming handles the first step, while MAI-Voice models address the final step, with the reasoning model Mai-Thinking-1 processing transcripts to determine the agent's actions.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at siliconangle.com →

More in AI

Robot study gives seniors say in era of AI

It’s no longer the stuff of sci-fi: AI sidekicks that simulate emotional support and guide routines. Robot companions are increasingly becoming part of everyday life for seniors, with certain U.S. […]

  • Seniors participate in AI robot study led by Clara Xie
  • BankBox simulates online banking for safe practice
  • Concern about AI's impact on older adults emphasized

How I chained three Apify Actors into a source-safe MCP agent

After making one business-registry Actor usable through the Apify MCP server, I tried the next obvious step: give an agent several company-research tools at once.

  • Author chains three Apify Actors into a single MCP agent
  • Actors used: US Business Entity Search, Domain Availability Checker, SEC EDGAR Company Filings
  • Output presents source-separated evidence bundle

More from Friday 2 October →