Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.