Urgent.News

What's breaking now, across thousands of outlets.

AI

Meta launches Muse Voice Transcribe, MSL's first real-time audio perception model, with streaming automatic speech recognition, trained with 70+ languages (Meta AI Research)

Experience Muse Voice Transcribe in real time … We're excited to introduce Muse Voice Transcribe, the first real …

Meta has unveiled Muse Voice Transcribe, a groundbreaking real-time audio perception model developed by Meta Superintelligence Labs. This innovation offers streaming automatic speech recognition, diarization for up to 20 speakers, and endpointing capabilities, all in a single, multilingual model.

The model outperforms in streaming speech-to-text and diarization benchmarks, ranking first on Artificial Analysis. It is an autoregressive multimodal model from the Muse Spark family, processing audio in 80ms chunks and transforming each into a single soft token. The model decides whether to continue listening or emit a text token at each audio chunk, with the ability to dynamically adjust the "delay" for each word based on difficulty.

Muse Voice Transcribe supports over 70 languages, with 25 extensively verified. It natively handles code-switching, a common feature in bilingual speech. It also introduces special tokens for diarization and endpointing, marking speaker switches and speech beginnings and ends. The model can transcribe long audio inputs and supports additional languages through the Meta Model API, Meta AI for Mac, and Muse Code.

Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at research.meta.ai →

More in AI

More from Tuesday 1 September →