Summary

On September 1, 2026, Meta Superintelligence Labs released Muse Voice Transcribe, its first real-time model combining streaming speech recognition, speaker diarization, and endpointing in one pass. It handles 20-plus speakers and audio over an hour, was trained on 70-plus languages with 25 validated at launch, and posts a 3.1% word error rate on Artificial Analysis's streaming benchmark, ahead of Cartesia, ElevenLabs, GPT, and Gemini live transcription.

What changed

Meta shipped Muse Voice Transcribe via the Meta Model API at $3 per 1,000 audio-minutes (about $0.18 per hour), unifying streaming ASR, speaker diarization, and endpointing without a separate post-processing step; it already powers dictation in Meta AI for Mac and Muse Code.

Why it matters

A single low-cost model that transcribes, separates speakers, and detects turn boundaries in real time simplifies voice pipelines for agents and apps that previously chained multiple services. Leading accuracy at aggressive pricing pressures specialized speech vendors and lowers the barrier to real-time voice interfaces.

Evidence excerpt

Muse Voice Transcribe combines streaming automatic speech recognition with speaker diarization and endpointing... word error rate is 3.1%, ahead of competitors like Cartesia Ink-2 (3.4%), ElevenLabs' Scribe v2 Realtime (3.6%)... available through the Meta Model API for $3 per 1,000 audio-minutes.

Sources