Microsoft Research has announced the release of the streaming speech recognition model "VibeVoice-ASR-Streaming-7B." This model features the ability to simultaneously perform speaker identification and content transcription in sync with audio input.

The model supports 10 languages. The supported languages are Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish.

Additionally, it includes a custom hotword feature. By pre-specifying words such as names or technical terminology, users can improve recognition accuracy within specific domains.

This project is released under the MIT License.


Source: microsoft/VibeVoice-ASR-Streaming-7B (HF: Microsoft, 2026-09-03)