Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue
Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop settles into. Two PersonaPlex-7B instances exchange audio tokens on a shared…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.