OmniVoice โมเดล TTS 600+ ภาษา โคลนเสียงจากคลิป 3 วินาที เปิดซอร์สฟรี
OmniVoice โมเดล TTS 600+ ภาษา โคลนเสียงจากคลิป 3 วินาที เปิดซอร์สฟรี ทีม k2-fsa เปิดตัว OmniVoice โมเดลแปลงข้อความเป็นเสียงพูด (text-to-speech) แบบ zero-shot ที่รองรับมากกว่า 600 ภาษา ซึ่งเป็นความครอบคลุมภาษาที่กว้างที่สุดในบรรดาโมเดล TTS แบบ zero-shot ที่มีอยู่ในปัจจุบัน โดยใช้สถาปัตยกรรม diffusion language model แบบใหม่ และรองรับทั้งการโคลนเสียง (voice cloning) และการออกแบบเสียง (voice design)…
OmniVoice TTS model has unveiled its ability to generate speech in over 600 languages through a three-second video clip. Developed by k2-fsa, the model utilizes a novel diffusion language model architecture and supports both voice cloning and voice design. Its key features include extensive language coverage, fast inference speed, and the capability to generate voices from short video snippets.
OmniVoice operates through Python API and command-line tools, supporting CUDA, Apple Silicon, and Intel Arc GPU. Users can clone voices from short video clips or create voices via instructions. The model is open-source and available on GitHub and Hugging Face, allowing developers to run it on local devices without incurring API charges.
For Thai users, OmniVoice offers the unique advantage of supporting the Thai language among the supported 600+ languages, enabling the creation of high-quality Thai speech on local devices without relying on cloud services. However, voice cloning involves ethical considerations, and users should verify the rights for using reference audio before implementation.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.