发表机构
School of Computer Science and Engineering; School of Design; South China University of Technology; Guangdong Engineering Center for Large Model and GenAI Technology; State Key Laboratory of Subtropical Building and Urban Science, Ministry of Education Key Laboratory of Big Data and Intelligent Robot; School of Computing and Information Systems(计算机科学与工程学院; 设计学院; 华南理工大学; 广东省大模型与生成式人工智能技术工程中心; 亚热带建筑科学与城市科学国家重点实验室,教育部大数据与智能机器人重点实验室; 计算与信息系统学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究二元对话运动生成问题,提出Learn2Chat框架,通过交互调制预训练单声道运动先验,引入单声道锚定运动分解等方案,能有效分离运动与交互效应,实验表明该框架性能领先且与模型无关,可有效实现可扩展对话运动生成。
AI 中文摘要
二元对话运动生成对于逼真的交互式数字人至关重要。现有方法通常在统一的二元生成器中对对话行为进行建模。然而,这种整体公式往往将自我语音驱动的运动与伙伴响应的社会反馈耦合在一起,使交互特定组件隐含,并未充分利用预训练单声道运动模型已经学到的语音-运动对应关系。我们提出Learn2Chat,一个统一框架,将二元运动建模为对预训练单声道运动先验的交互调制。此设计将内在语音驱动的运动与社会交互效应分开,并实现更结构化的交互建模。具体而言,我们引入了一种单声道锚定运动分解方案,利用从单声道数据中学到的语义运动流形,将音频驱动的运动动力学与交互诱导的调制分开,从二元序列中产生干净的交互表示。在此表示空间之上,一个交叉注意力交互潜在预测模块通过跨分支注意力和交互对齐将配对的语音信号映射到交互潜在。在推理过程中,预测的交互潜在调制规范的单声道运动,以数据高效的方式生成连贯和同步的二元行为。在DualTalk基准上的大量实验表明,Learn2Chat在定量指标和感知评估方面都取得了领先性能。此外,该框架与模型无关,并与各种预训练的单声道运动主干无缝集成,突出了先验重用和交互适应对可扩展对话运动生成的有效性。项目页面上有更多视觉结果。
英文摘要
Dyadic conversational motion generation is essential for realistic interactive digital humans. Existing approaches typically model conversational behaviors within unified dyadic generators. However, such holistic formulations tend to couple self-speech-driven motion with partner-responsive social feedback, leaving the interaction-specific component implicit and underutilizing the speech-motion correspondence already learned by pretrained monologic motion models. We propose Learn2Chat, a unified framework that models dyadic motion as interaction modulation over pretrained monologic motion priors. This design separates intrinsic speech-driven motion from social interaction effects and enables more structured interaction modeling. Specifically, we introduce a Monologic-Anchored Motion Factorization scheme that leverages the semantic motion manifold learned from monologic data to disentangle audio-driven motion dynamics from interaction-induced modulation, yielding clean interaction representations from dyadic sequences. On top of this representation space, a Cross-Attentive Interaction Latent Prediction module maps paired speech signals to interaction latents through cross-branch attention and interaction alignment. During inference, the predicted interaction latents modulate canonical monologic motion to generate coherent and synchronized dyadic behaviors in a data-efficient manner. Extensive experiments on the DualTalk benchmark demonstrate that Learn2Chat achieves state-of-the-art performance across both quantitative metrics and perceptual evaluations. Moreover, the framework is model-agnostic and seamlessly integrates with diverse pretrained monologic motion backbones, highlighting the effectiveness of prior reuse and interaction adaptation for scalable conversational motion generation. More visual results are available on the project page.