Urgent.News

What's breaking now, across thousands of outlets.

Tech

Meta Launches Muse Voice Transcribe for Real-Time Multilingual Speech

Meta’s Muse Voice Transcribe delivers real-time multilingual speech recognition, speaker labeling and adaptive latency for developers and Mac users. The post Meta Launches Muse Voice Transcribe for Real-Time Multilingual Speech appeared first on TechRepublic .

Meta has unveiled Muse Voice Transcribe, a real-time multilingual speech recognition technology that marks its first real-time audio perception model. Unlike traditional transcription systems that process recordings after the fact, Muse delivers text continuously while identifying speakers and detecting speech boundaries. The model supports up to 20 speakers and recognizes when speech switches between languages during a conversation.

Trained across over 70 languages, with 25 extensively validated for the initial release, Muse is particularly useful in multilingual markets like India, where conversations may shift between English and regional languages such as Hindi, Tamil, Telugu, Kannada, and Malayalam. One of Muse's key technical features is its adaptive latency approach, which processes audio in 80-millisecond chunks and decides how long to listen before committing to each word.

This allows easy words to be transcribed quickly while more challenging ones receive additional audio context before the model makes a decision. This "adaptive delay" is trained using reinforcement learning to balance speed and accuracy, aiming to reach the Pareto front for speed and accuracy. For consumers, the most immediate application is in Meta AI for Mac, where Muse Voice Transcribe now powers dictation features.

Developers can access the model through Meta’s Model API at $3 per 1,000 audio minutes, roughly 18 cents per hour. While Muse offers significant benefits, particularly in reducing the need for separate speech-processing systems, developers should test its performance with their own accents, background noise, terminology, and language combinations, as performance may vary across languages and real-world recordings.

Written by urgent.news from TechRepublic's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techrepublic.com →

More in Tech

More from Wednesday 2 September →