On September 1 (local time), Meta announced "Muse Voice Transcribe," the first real-time speech recognition model developed by Meta Superintelligence Labs (MSL). The model is provided through the "Meta Model API" and is priced at $0.18 per hour.
Model ReleasesPricing & LimitsMetaMuse Voice Transcribe
Meta Announces Muse Voice Transcribe, Its First Real-Time Speech Recognition Model
This article is a translation. Read the Japanese original
Using a single model, it processes streaming ASR (Automatic Speech Recognition) as well as speaker diarization and endpointing. Meta stated that the model can distinguish between more than 20 speakers and can handle audio exceeding one hour without requiring post-processing.
The model was trained on over 70 languages, with 25 languages, including Japanese, verified at the time of initial release. It natively handles "code-switching," where languages switch within a single sentence, and recognition accuracy can be improved by specifying language or contextual biases.
Technically, it is positioned as an autoregressive multimodal model within the Muse Spark family. It processes audio in 80-millisecond chunks, with the model itself deciding whether to continue listening to the audio or output text. Through reinforcement learning, it introduces "adaptive delay" to improve the tradeoff between accuracy and latency.
Regarding performance, Meta stated that the model ranked first in the streaming speech recognition rankings by Artificial Analysis. According to data from Artificial Analysis, the Word Error Rate (WER) for finalized transcriptions is 3.1%, and the latency from the end of speech is 0.16 seconds, surpassing models from Cartesia and ElevenLabs.
In addition to the API for developers, the model is available in the Mac version of Meta AI and in Muse Code, a coding product. Notably, there is no mention of releasing the model weights in the official blog, meaning availability is currently limited to the API and Meta's own applications.
Source: Meta、初のリアルタイム音声認識モデル「Muse Voice Transcribe」 20人超の話者識別と多言語混在に対応 (ITmedia AI+, 2026-09-02)