English

Model ReleasesMicrosoftMAI-Transcribe-2

Microsoft Releases MAI-Transcribe-2 Speech Recognition Model, Claiming Higher Speed and Accuracy Than Competitors

This article is a translation. Read the Japanese original

Microsoft has released "MAI-Transcribe-2," a speech recognition model capable of automatically detecting and transcribing 60 different languages.

According to Microsoft, the model ranked first in the "FLEURS" benchmark with an average word error rate of 5.2%, surpassing competing models such as "Gemini 3.5 Transcribe," which was announced one week prior.

Verification by Artificial Analysis showed that the processing speed is 10 times faster than OpenAI's "GPT-Transcribe," 7 times faster than ElevenLabs' "Scribe v2," and 5 times faster than Google's "Gemini 3.5 Transcribe." Simultaneously, it is reported to demonstrate higher accuracy.

In terms of functionality, the model excels at speaker identification, word timing synchronization, and the recognition of technical terminology. It also includes a setting that allows users to choose whether to retain or remove fillers, such as "uh" and "um."

Regarding cost, the price has been set at $0.1 per hour for a limited time until the end of 2026. It has been pointed out that this is more affordable than Scribe V2 and Smallest AI Pulse Pro.


Source: Microsoftが文字起こしAI「MAI-Transcribe-2」をリリース、Gemini 3.5 Transcribeより5倍高速&GPT-Transcribeより10倍高速 (GIGAZINE, 2026-09-04)