GDPO-Listener: 通过自动回归流匹配和组奖励解耦策略优化实现表达性交互头部生成
GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization
浏览论文内容
中文总结 AI 辅助
本文提出GDPO-Listener框架,通过自动回归流匹配和组奖励解耦策略优化,实现高表达性的说话和倾听动作生成,提升了长期运动变化、视觉表达性和语义可控性。
中文摘要 AI 辅助
生成逼真三维头部运动对于双人交互是虚拟人合成中的重大挑战。尽管近期方法在说话头部方面取得显著成果,但监听动作常出现'回归均值'问题,导致动作塌陷为静态面孔,并缺乏参数空间以生成复杂非言语动作。本文提出GDPO-Listener,一种新的框架,实现高度表达性的说话和倾听动作生成。首先,我们引入了自动回归流匹配架构,实现稳定的监督学习。其次,为克服运动僵硬,我们应用组奖励解耦策略优化(GDPO)。通过在不同的FLAME参数组间隔离奖励归一化,GDPO明确激励高方差的表达生成。最后,我们实现了显式的语义文本控制以实现定制化响应。在Seamless Interaction和DualTalk数据集上的广泛评估显示,其在长期运动变化、视觉表达性和语义可控性方面优于现有基线。
英文摘要
Generating realistic 3D head motion for dyadic interactions is a significant challenge in virtual human synthesis. While recent methods achieve impressive results with speaking heads, they frequently suffer from the `Regression-to-the-Mean' problem in listener motions, collapsing into static faces, and lack the parameter space for complex nonverbal motions. In this paper, we propose GDPO-Listener, a novel framework that achieves highly expressive speaking and listening motion generation. First, we introduce an Auto-Regressive Flow Matching architecture enabling stable supervised learning. Second, to overcome kinematic stillness, we apply the Group reward-Decoupled Policy Optimization (GDPO). By isolating reward normalization across distinct FLAME parameter groups, GDPO explicitly incentivizes high variance expressive generations. Finally, we enable explicit semantic text control for customizable responses. Extensive evaluations across the Seamless Interaction and DualTalk datasets demonstrate superior performance compared to existing baselines on long-term kinematic variance, visual expressivity and semantic controllability.
发表机构
- University of Southern California, Institute for Creative Technologies(南加州大学创意技术研究所)
机构由 AI 辅助整理,请以论文原文为准。