Automating the Selection of Natural, High-Quality Single-Speaker Anchors
๐ Originally published (in Japanese) at forge.workstyle.tech . When developing a voice conversion app, you often find that the quality of the conversion depends more on the choice of the target voice (anchor) than on the conversion model itself. In our voice conversion app, we used 4 to 12 seconds of recordings from voice actors or narrators as anchors. Initially, we relied on "what sounded goodโฆ
Developing a voice conversion app often involves selecting an anchor from the recordings of voice actors or narrators. Initially, selecting these anchor segments was done based on subjective judgment of what sounded good. However, this method was both time-consuming and inconsistent due to variations in mood and individual opinions.
Moreover, the quality of the recordings varied, with some containing overlapping voices, excessive intonation, or background noise. To overcome these challenges, the article details a three-stage filter system that automatically selects suitable anchor segments. The goal was to choose segments that are single-speaker, natural, and of high quality.
The system improved the minimum PESQ score of the anchor set from 1.24 to 2.31 and ultimately compiled over 70 anchors.
Written by urgent.news from Dev.to's reporting โ not their text. Machine-written โ may contain errors; check the original before relying on it.