A free, self-hosted AI studio with 200+ unfiltered
Open-source multi-speaker TTS for conversational audio and podcast generation.
A Frontier Open-Source Text-to-Speech Model
VibeVoice is a novel framework by Microsoft Research Asia's GeneralAI Group designed for generating expressive, long-form, multi-speaker conversational audio from text. It addresses key challenges in traditional TTS, including scalability, speaker consistency, and natural turn-taking, making it ideal for producing podcast-style content and other dialogue-driven audio.
VibeVoice is completely free and open-source. You can clone the repository, download the models, and run them on your own infrastructure at no cost. Be sure to review the specific open-source license for any redistribution requirements.
Summarized from the official site: https://microsoft.github.io/VibeVoice
Yes, VibeVoice is completely free and open-source. You can clone the repository, download the models, and run them on your own infrastructure at no cost.
Yes, VibeVoice is fully open-source. Its code, model weights, and tools are available on GitHub and Hugging Face, though you should review the specific license for redistribution requirements.
VibeVoice is an open-source text-to-speech framework developed by Microsoft Research Asia. It is designed to generate expressive, long-form, multi-speaker conversational audio from text.
Yes, VibeVoice supports multi-speaker generation. It creates audio with multiple distinct speakers, each maintaining a consistent voice throughout long sequences.
You can use VibeVoice to automatically generate multi-speaker podcasts, narrate audiobooks with distinct character voices, and power conversational AI voice interfaces.
A free, self-hosted AI studio with 200+ unfiltered
Optimize and secure AI coding agents with this ope
Generate highly realistic AI images from text with
Cloud-based text-to-speech service with 47 natural voices.