Skip to main content

Meta's New AI Model Transcribes 20 Speakers at Once, Costs Just $3 per 1,000 Minutes

Meta has just dropped its first real-time audio perception model, and it's a game-changer for anyone who's ever struggled to keep up with a fast-paced conversation or a chaotic meeting. Meet Muse Voice Transcribe, a model that can transcribe speech as it happens, separate up to 20 speakers, and even understand when someone pauses or stops talking. And the best part? It's surprisingly affordable at just $3 per 1,000 minutes of audio—that's about $0.18 per hour.

Speak, and It Types—Automatically Separating 20+ Speakers

Imagine you're in a meeting with a dozen people, all talking over each other. Traditional transcription tools would give you a garbled mess. But Muse Voice Transcribe is different. It uses streaming automatic speech recognition, which means it transcribes as people speak, not after the fact. No more waiting for the recording to finish before you get your text. Instead, you get real-time output, segment by segment, making it perfect for live subtitles, meeting notes, or even call transcriptions.

What's truly impressive is its ability to separate speakers. Even with more than 20 people talking, the model can tell who's who. It also detects when a speaker stops, automatically marking the end of their turn. This is a huge leap forward for transcription accuracy in group settings.

Language-wise, Muse Voice Transcribe is a polyglot. It supports over 70 languages, with 25 verified at launch. It can handle audio longer than an hour and even switch languages mid-sentence or between sentences—no need to restart or reconfigure. Developers can also fine-tune recognition for specific jargon or scenarios by using language, keyword, and context bias.

Adaptive Delay: The Secret to Balancing Speed and Accuracy

You might be wondering: how does it manage to be both fast and accurate? The answer lies in a clever mechanism called "adaptive delay." Instead of making a one-size-fits-all trade-off, the system decides for each word how much additional audio it needs before outputting the result. For clear, easy-to-recognize speech, it spits out the transcription almost instantly. But for words that are mumbled, full of technical terms, or in complex contexts, it waits a bit longer, gathering more context to ensure accuracy. It's like a human listener who nods along quickly when things are clear but leans in and asks for clarification when something's ambiguous.

According to Meta, as of September 1, 2026, Muse Voice Transcribe has already claimed the top spot on the Artificial Analysis Streaming Speech-to-Text Leaderboard. That's a strong endorsement of its capabilities.

What This Means for Developers and Users

For developers, this opens up a world of possibilities. Real-time transcription can power live captioning for webinars, assist in customer service calls, or even help journalists transcribe interviews on the fly. The pricing is competitive, making it accessible for startups and small businesses. And with the ability to handle multiple speakers and languages, it's a versatile tool for global teams.

But it's not just about the tech specs. This model feels like a step toward more natural human-computer interaction. Instead of talking to a machine that waits for you to finish, you're having a conversation where the machine keeps up with you. It's like having a super-efficient assistant who never misses a word.

Key Points

  • Real-time transcription: Streams text as you speak, reducing latency.
  • Speaker separation: Handles up to 20+ speakers, identifying who said what.
  • Language support: Over 70 languages, with 25 verified at launch.
  • Adaptive delay: Balances speed and accuracy dynamically.
  • Affordable pricing: $3 per 1,000 minutes of audio.
  • Top-ranked: #1 on the Artificial Analysis Streaming Speech-to-Text Leaderboard.

Image

Image

Muse Voice Transcribe is more than just a new model—it's a glimpse into a future where machines listen and understand in real time, making our interactions smoother and more efficient. Whether you're a developer looking to integrate cutting-edge speech recognition or a business owner tired of messy meeting notes, this could be the tool you've been waiting for.