Urgent.News

What's breaking now, across thousands of outlets.

Tech

Where to split a sentence for streaming TTS is decided by one number

I build a desktop app where you talk to a VRM avatar and it answers out loud. Someone told me the gap between speaking and hearing a reply was too long. I measured it and fixed it. Then I found out my benchmark had been running on a machine with a full disk and a load average in the low hundreds, which made the number 4x worse than reality , and the design decision I derived from it was wrong for…

The time it takes for text-to-speech (TTS) to split a sentence for streaming is determined by a single number called the real-time factor (RTF). A developer discovered this by measuring the gap between speaking and hearing a reply in their desktop app that used VRM avatars. However, their benchmark was run on a machine with a full disk and high load, causing the RTF to be four times worse than the actual performance. This led to an incorrect design decision for a week before they realized their mistake.

Key points:

1. The RTF is the crucial factor in determining where to split a sentence for streaming TTS.

2. Store the formula for calculating the RTF rather than the specific number it produces, as different machines will yield different results.

3. Seconds are not the right unit for comparing TTS engines; instead, focus on the ratio of synthesis time to audio length.

4. Longer sentences take longer to synthesize, but the important metric is the ratio against the length of the audio produced.

5. The split point can be calculated using the formula: p = r / (1 + r), where p is the fraction of the sentence in the first chunk, and r is the RTF.

6. At an RTF of 0.25, cutting the first chunk at 20% of the sentence length results in no gap. At an RTF of 1.0, you must wait until halfway through the sentence before starting.

7. Check that the machine is idle before benchmarking and analyze the RTF instead of relying on wall clock time.

8. Implement a queue for playback to ensure no gaps occur when synthesizing the first and subsequent chunks of the sentence.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Tuesday 11 August →