发表机构
Stony Brook University; Atmanity Inc(石溪大学; Atmanity公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GLARE,一种基于流匹配Transformer的音频驱动聆听头部生成模型,并构建了带细粒度反应标注的数据集及反应导向评估协议,以提升双人对话中聆听者反应的恰当性。
AI 中文摘要
尽管说话头部生成技术已迅速发展,但在双人对话中生成自然的聆听者行为(即知道何时反应、如何反应以及以何种类型的反应进行回应)仍未得到充分探索。现有的双人对话数据集缺乏细粒度的聆听者反应标注,而从说话头部和视频生成中继承的现行评估指标衡量的是视觉真实感,而非聆听者是否做出了恰当的反应。我们针对这三个方面弥补了这些空白。首先,我们构建了一个专门针对聆听头部的数据集,该数据集基于RealTalk和Seamless Interaction构建,包含约147小时的配对说话者-聆听者视频,以及涵盖六种类别(点头、摇头、微笑、大笑、皱眉和惊讶)的64,557个事件级反应标注。其次,我们引入了一个基于流匹配Transformer的音频驱动基线模型,即GLARE,其韵律条件由Qwen2-Audio提取,并采用时间反应损失来显式监督逐帧反应。第三,我们提出了一种面向反应的评估协议,该协议联合衡量反应发生情况(R-F1)、时间对齐(R-tIoU)、非对称时间偏差(R-ATD)和反应区域视觉质量(R-FID),从而提供比仅基于视觉质量的指标更具行为学基础的评估。实验结果表明,在视觉保真度和反应级指标方面,该方法均优于先前的聆听头部方法,这表明反应感知的数据、建模和评估对于自然的聆听行为至关重要。
英文摘要
While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.
CommentsAccepted in NeurIPS 2026. Project page: https://github.com/lzk901372/glare