AInspiro
Tool Reviews

Meta Dropped a Price Bomb: It Tells 20 Speakers Apart, and 1,000 Minutes Costs $3

AInspiro Editorial·
This article was created with AI assistance.

A 20-person meeting ends, and the notes are a mess

Sales, engineering and marketing all talk at once. A remote colleague jumps in. When the meeting ends, the person writing the summary faces a tangle of who said what. This is not a joke. It is the real, repeated cost of hundreds of millions of meetings every day.

On September 1, Meta answered with something unexpected: not pricier hardware, but an API that charges 3 US dollars per 1,000 audio-minutes, about 0.18 US dollars an hour, roughly 1.2 yuan.

What it is: three systems collapsed into one model

This is Meta Superintelligence Labs' first real-time audio perception model, Muse Voice Transcribe. Most production voice stacks stitch three pieces together: one transcribes, one separates speakers, one decides when a sentence ends. Every hand-off adds latency and a point of failure.

Muse inverts that. Streaming speech recognition, speaker diarization for more than 20 voices, and endpoint detection all sit inside one autoregressive model. Audio arrives in 80-millisecond chunks, each turned into a single soft token, and after every chunk the model decides: keep listening, or emit the text. This adaptive delay is trained with reinforcement learning. Hard words get more listening time, clear words commit early, instead of one fixed wait applied to every sentence.

The scores and the price

Language coverage: trained on 70-plus languages, 25 validated at launch. A single audio clip can run over an hour and supports code-switching within or between sentences. Developers can bias recognition with keywords and context to lift specific terms.

On benchmarks it ranks first on the Artificial Analysis streaming speech-to-text leaderboard. As of September 1, 2026, its final word-error rate is 3.1%, ahead of Cartesia at 3.4% and ElevenLabs at 3.6%. Its diarization error rate is 17.5%, also beating the field in the same comparison.

On price it lands in the same tier as OpenAI's GPT-4o Mini Transcribe, but streaming speaker separation is its real differentiator. The model is live, powering dictation in the Mac app and in Muse Code, and available through the Meta Model API as muse-voice-transcribe-1.0.

The boundaries, stated plainly

It is not the cheapest on the market. Soniox publishes a lower equivalent rate. Its speaker count is not the highest either: Speechmatics documents up to 100 speakers and Amazon up to 30, both above its 20-plus. Meta's pitch is not leading on any single number. It is bundling high-capacity diarization, low latency and aggressive pricing into one model.

Another honest point: that 3.1% word-error rate is a benchmark figure. Your accents, your industry terms, your noisy meeting room need a real-audio test before you trust it. The 80-millisecond streaming delay is fine for live captions but must be stress-tested for hard-realtime use. Do not take the marketing line at face value.

What this means for you

If you build meeting, call-center or interview transcription products, this API belongs on your test list. For the same money, you used to get transcription without speaker labels. Now you get per-speaker, timestamped output, and that can lift the product a full step.

But do not flip everything over at once. First test three things with your own recordings: Chinese names and places, separation accuracy when people talk over each other, and whether 80 milliseconds is truly enough in your scenario. Pass those three, then talk about replacement.

Where it actually differs from rivals

Put Muse next to OpenAI, ElevenLabs and Google and the raw transcription accuracy sits in a narrow 3% to 4% WER band. The real watershed is bundling speaker separation, streaming and low price into one model. Before this you stacked three services: one to transcribe, one for diarization, one for endpointing, each billing separately with its own latency. Now it is one bill.

There is a quiet plus for Chinese users: Meta's training corpus covers 70-plus languages, and Chinese is baseline, not an edge case. But accents and proper nouns still need your own test. Do not go live on the leaderboard ranking alone.

A practical caveat on the 3 dollars

The 3 dollars per 1,000 minutes is a list price for processed audio, not a ceiling on your real cost. Long silences, overlapping speech and repeated retries still consume minutes. For a meeting-heavy team, model the spend on your actual recording patterns before committing, because the number that matters is cost per useful transcribed hour, not cost per raw minute.