ChengRang

VibeVoice

AI Audio & Music Free

Microsoft open-source voice AI model family with TTS and ASR, supporting 60-min long audio recognition and 90-min multi-speaker synthesis

MicrosoftOpen SourceSpeech RecognitionTTSASR
Visit VibeVoice

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

VibeVoice is a cutting-edge family of speech AI models developed and open-sourced by Microsoft Research, covering two core capabilities: Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The research paper was published in August 2024, and the project was officially open-sourced on GitHub in March 2026, garnering over 45K stars by the end of April. Its core innovation is a 7.5Hz ultra-low frame rate tokenizer, achieving a 3200x audio compression ratio. Combined with a hybrid LLM + diffusion head architecture, it achieves industry-leading performance in long audio processing and multi-speaker synthesis.

Key Features

Use Cases

Pros

Pricing

Completely free and open-source (MIT license), with code and model weights hosted on GitHub. It can also be accessed as a cloud service through the model catalog on the Microsoft Foundry platform.

Summary

VibeVoice is one of the most comprehensive solutions in the current open-source speech AI field, far surpassing traditional solutions in both long audio processing and multi-speaker synthesis. It is suitable for developers and enterprises that require high-quality speech recognition/synthesis and value data privacy.

Version History

Category
AI Audio & Music
Pricing
Free
Tags
Microsoft · Open Source · Speech Recognition
Website

Related Tools