Urgent.News

What's breaking now, across thousands of outlets.

Tech

Meta’s Muse Voice Transcribe 70+ Languages and 20+ Speakers Simultaneously

Meta has introduced Muse Voice Transcribe, its first real-time audio model designed for live dictation and transcription across multiple speakers … Read More The post Meta’s Muse Voice Transcribe 70+ Languages and 20+ Speakers Simultaneously appeared first on ProPakistani .

Meta’s Muse Voice Transcribe 70+ Languages and 20+ Speakers Simultaneously

Meta has unveiled Muse Voice Transcribe, a groundbreaking real-time audio model capable of simultaneous live dictation and transcription of over 70 languages and 20 speakers. This innovative technology is the first of its kind, designed to handle complex scenarios such as multiple speakers, language switching, and code-switching within a single conversation.

During a demonstration, Meta CEO Mark Zuckerberg showcased the model's ability to automatically identify different speakers and seamlessly transition between languages during a conversation. The model employs an adaptive delay mechanism to enhance transcription accuracy, carefully processing simpler words while allowing longer for more challenging words. Initially trained on more than 70 languages, 25 languages were validated at launch, allowing users to take advantage of Muse Voice Transcribe's capabilities right away.

Muse Voice Transcribe is designed to tackle noisy and intricate real-world audio, mid-sentence language switching, and long sessions with multiple speakers. Mark Zuckerberg shared the demonstration after returning to the social media platform X after roughly three years of not posting. This first real-time audio perception model is a product of Meta's Superintelligence Lab, following recent releases that include an open-weight coding agent and the Meta AI Mac app.

Muse Voice Transcribe can be accessed through Meta AI's Mac app, which can provide voice features to other applications. Developers can integrate the model into their services via Muse Code and Meta's Model API, with pricing set at $3 per 1,000 minutes of audio. A demonstration version is also available on Meta's research blog.

Written by urgent.news from ProPakistani's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at propakistani.pk →

More in Tech

What is harness engineering and why should I care?

How do you ship a software product with 0 lines of manually-written code? A friend asked me this today, and I realized I didn't have a simple answer. So I dug deeper.

  • Harness engineering designs environment for AI agents to operate safely and reliably.
  • A well-engineered harness ensures AI agent runs within right parameters and prevents disruptions.

More from Wednesday 2 September →