Microsoft VibeVoice
VibeVoice TTS
Long-form multi-speaker synthesis
Windows only
Generation
Up to 90 min
Speakers
Up to 4
Platform
Windows only
OpenVox fit
Dialogue + long-form
Model overview
Built into OpenVox for private local voice generation.
VibeVoice TTS is a long-form, multi-speaker text-to-speech model designed for stable conversational audio and extended scripted generation.
VibeVoice can synthesize long-form speech of up to 90 minutes in a single generation workflow.
It supports up to four distinct speakers for conversations, podcasts, scripts, and narrated dialogue.
In OpenVox, VibeVoice TTS is available on Windows only and runs locally after its model files are downloaded.
OpenVox does not provide celebrity voice models and does not permit cloning, impersonating, or commercially using any person's voice without proper rights or consent.
Open-source model: Microsoft VibeVoice
View on Hugging Face