Where to split a sentence for streaming TTS is decided by one number
I build a desktop app where you talk to a VRM avatar and it answers out loud. Someone told me the gap between speaking and hearing a reply was too long. I measured it and fixed it. Then I found out my benchmark had been running on a machine with a full disk and a load average in the low hundreds, which made the number 4x worse than reality , and the design decision I derived from it was wrong for…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.