发表机构
Hefei University of Technology; EPIC Lab, Shanghai Jiao Tong University; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center; SAI, Shanghai Jiao Tong University; University of Science and Technology of China; United Arab Emirates University; Anhui University(合肥工业大学; 上海交通大学EPIC实验室; 合肥综合性国家科学中心人工智能研究院; 上海交通大学SAI; 中国科学技术大学; 阿联酋大学; 安徽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EvolvingAvatar利用测试时训练和双人上下文预测,在交互中自适应生成三维头部动作,并通过InterHead-Bench基准验证,显著改善对话运动统计。
AI 中文摘要
交互式三维头部生成需要协调的说话和聆听动作,以响应不断演变的对话。现有生成器将传入的观察作为上下文,但保持参数固定,未将对话模式用作学习信号。我们提出EvolvingAvatar,一种因果生成器,在交互期间利用测试时训练来适应用户面部视频和双人音频。其双人上下文预测目标提供来自音视频上下文的自监督学习信号,而无需测试时的目标运动标签。持久快速权重在每次对话内累积这些更新以引导运动生成,而瞬时下颌适应响应当前音视频上下文。预测的语音活动控制持久适应如何引导运动。我们还引入InterHead-Bench,一个由单视角和双视角对话视频构建的统一455.95小时基准。实验表明,与强基线相比,对话运动统计得到改善。在最困难的分布外分割中,随着对话展开,生成得到改进,从第一个时间间隔起,记录的用户-头像表情统计失配最多减少11.1%。
英文摘要
Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAvatar, a causal generator that uses test-time training to adapt to user face video and dyadic audio during interaction. Its dyadic context prediction objective provides a self-supervised learning signal from audiovisual context without target motion labels at test time. Persistent fast weights accumulate these updates within each conversation to guide motion generation, while transient jaw adaptation responds to current audiovisual context. Predicted speech activity controls how persistent adaptation guides motion. We also introduce InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos. Experiments show improved conversational motion statistics over strong baselines. On the hardest out-of-distribution split, generation improves as conversations unfold, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.
CommentsProject Page: https://blog.evolving-avatar.com