microsoft/VibeVoice-ASR-Streaming-1.5B

Microsoft releases VibeVoice-ASR-Streaming-1.5B, a unified streaming speech recognition model that attributes speakers and transcribes content simultaneously. The model supports ten languages and allows customized hotwords for domain-specific accuracy. Developers can access the code and usage instructions via the linked GitHub repository under an MIT license.

README

microsoft/VibeVoice-ASR-Streaming-1.5B View on Hugging Face

Loading the README from Hugging Face…