Microsoft Research has released the streaming ASR model "VibeVoice-ASR-Streaming-1.5B."

This model features "Speaker-Attributed Transcription," which continuously transcribes who said what simultaneously as audio is received.

Users can specify custom hotwords, such as names and technical terminology, making it possible to improve recognition accuracy for domain-specific content.

The model supports 10 languages: Japanese, English, Chinese, French, German, Italian, Korean, Portuguese, Russian, and Spanish.

This project is released under the MIT License.


Source: microsoft/VibeVoice-ASR-Streaming-1.5B (HF: Microsoft, 2026-09-03)