Microsoft targets ultra-realistic voice agents with its first streaming transcription model
Microsoft Corp. today expanded its MAI artificial intelligence model family with its first streaming transcription model, debuting alongside two others focused on text-to-speech. They’re designed for developers who want to build voice agents that can listen to people’s voices and reply instantly, similar to how humans talk to one another. The most consequential of the […] The post Microsoft…
Microsoft has launched its first streaming transcription model, MAI-Transcribe-2-Streaming, as part of its expanding MAI artificial intelligence model family. This model is designed to enable developers to create voice agents that can listen to human speech and respond instantly, similar to natural conversation. The model accepts human speech via a WebSocket, continuously updating the transcript as the person keeps talking, and confirms when the transcript is final.
This allows applications to display live captions or process a user's request before they have finished speaking.
MAI-Transcribe-2-Streaming is priced at $0.54 per audio hour, supports over 60 languages, and can automatically detect the language being spoken. It delivers first transcript hypotheses within an average of 320 milliseconds, though response speed may vary depending on network connection and the AI system generating the response.
The company also introduced two additional models focused on text-to-speech generation: MAI-Voice-2.1, which produces more expressive and high-fidelity outputs, and MAI-Voice-2.1-Flash, which offers faster response times and lower costs. Both models support 23 languages.
These new models demonstrate Microsoft's accelerated transition away from model providers like OpenAI and Anthropic, despite being a major investor in both. Microsoft AI Chief Executive Mustafa Suleyman has been concerned about the costs associated with these powerful frontier models and has instructed AI researchers to focus on developing the MAI model family to reduce costs and power Copilot agents in platforms like Excel and Outlook.
With this launch, Microsoft now has the necessary components to build powerful voice AI agents that can converse with users in a natural, humanlike manner. Voice agents require the ability to recognize and understand speech, decide on actions based on what was said, and generate audible responses. MAI-Transcribe-2-Streaming handles the first step, while MAI-Voice models address the final step, with the reasoning model Mai-Thinking-1 processing transcripts to determine the agent's actions.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.