Meta Launches Muse Voice Transcribe for Real-Time Audio and Multi-Speaker Processing
Tech

Meta Launches Muse Voice Transcribe for Real-Time Audio and Multi-Speaker Processing

TechNews Editorial
TechNews EditorialSep 2, 2026 · 1 min read
Share

Meta has released Muse Voice Transcribe, its first real-time audio model. The system provides dictation and transcription for more than 20 speakers while seamlessly managing multiple languages at the same time.

Mark Zuckerberg shared a demonstration video highlighting the software. The clip shows the transcription automatically separating different voices and switching between languages on the fly. The system also detects code-switching, which involves transcribing sentences that mix vocabulary from multiple languages.

The technology comes from the Meta Superintelligence Lab. The lab describes the software as state-of-the-art in streaming speech-to-text. It natively executes speaker diarization and endpointing inside a single model.

Zuckerberg explained the mechanics behind the adaptive delay feature. The model decides when to listen by waiting longer on difficult words and committing faster on simple ones. This method predicts each token to improve accuracy.

The software handles messy real audio conditions. It trained across more than 70 languages, with 25 of them validated at launch. It processes mid-sentence code-switching and manages hour-long sessions featuring over 20 distinct speakers.

This rollout follows Google Gemini 3.5 Transcribe by less than a week. The rival Google audio model features similar capabilities. Google plans to integrate its model into Android and Chrome, but Meta has not clarified its broader integration plans for flagship services.

Users can currently test the new capabilities through the Meta AI Mac app. The desktop application powers voice features across other services, meaning Muse Voice Transcribe now drives dictation functions in those programs.

Developers can access the model through Muse Code and the Meta Model API. Pricing is set at $3 for every 1,000 audio minutes. A public demo version is available on the Meta research blog.

Muse Voice Transcribe represents the newest software from the Meta Superintelligence Lab. The division recently released a dedicated coding agent, an open-weight model, and the Meta AI Mac app.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Related Stories