OmniVoice TTS model 600+ languages, voice cloning from 3-second clips, free open source.
OmniVoice TTS model 600+ languages, clone voices from 3-second clips, free open source. The k2-fsa team launched OmniVoice, a zero-shot text-to-speech model that supports more than 600 languages, which is the widest language coverage among existing zero-shot TTS models. It uses a new diffusion language model architecture and supports both voice cloning and voice design.
The k2-fsa team has introduced OmniVoice, a zero-shot text-to-speech model that supports over 600 languages. This model uses a new diffusion language model architecture and allows for voice cloning and design. OmniVoice can clone a voice from a 3-15 second audio clip and generate speech in a specific accent or style. The model is open-source and available on GitHub and Hugging Face.
Written by urgent.news from Dev.to's report — not a translation of it. Machine-written — may contain errors; check the original before relying on it.