先听后说:基于听者面部反应的对话语音生成响应规划
Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
浏览论文内容
中文总结 AI 辅助
本文提出ReACT-TTS两阶段框架,利用听者说话前1秒的面部反应序列规划对话语音的情绪和韵律,实验表明时间条件化优于文本条件化,且听者反应动态可作为对话响应规划的有效补充线索。
中文摘要 AI 辅助
对话语音依赖于对话上下文和听者紧接着的前序行为。我们提出ReACT-TTS,一个两阶段框架,利用说话前1秒的听者面部序列,在语音实现之前规划下一句话的情绪和韵律。在严格的二元MELD协议下,跨10个随机种子,时间条件化相比仅文本条件化获得了更高的平均宏F1分数和VAD一致性,而准确率基本保持不变。消融实验表明,在视觉变体中,时间建模表现最佳,且显式的早期到晚期差异是不必要的;正确的听者反应平均上也优于循环不匹配。在一项由20位语音研究人员参与的上下文适当性研究中,76%的判断偏好时间条件化,9%偏好仅文本,15%报告无偏好。我们进一步将预测的响应风格连接到Grad-TTS骨干网络,实现端到端的语音生成。总体而言,结果支持将听者前序反应动态作为对话响应规划的补充线索。源代码可在该https URL获取。
英文摘要
Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at https://github.com/CYJ1/ReACT-TTS_public.
发表机构
- Sogang University(西江大学)
机构由 AI 辅助整理,请以论文原文为准。