arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STEER:可控二元头部头像

STEER: Steerable Dyadic Head Avatars

Kartik Teotia, Helge Rhodin, Hyeongwoo Kim, Marc Habermann, Christian Theobalt

arXiv 2607.23840首次发表:更新:

发表机构

Max Planck Institute for Informatics; Saarland Informatics Campus(马克斯·普朗克信息研究所; 萨尔兰信息学园区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对语音驱动面部动画难以明确控制非语言渠道问题,提出STEER方法,将对话行为分解为多方面控制,通过特定管道获取行为伪标签,利用变压器学习目标运动,嵌入头像管道实现可控动画,在多方面表现优异。

AI 中文摘要

面部运动和表情在面对面交流中至关重要,能传达轮流发言、注意力、同意和参与等信息。尽管语音驱动的面部动画在唇同步和音频条件运动生成方面取得了很大进展,但大多数方法将对话行为视为音频的附带产物,或仅提供粗略的序列级情感控制。因此,关键的非语言渠道,如目光接触和回避、有节奏的头部运动和情感,仍然难以明确控制。我们提出了STEER,一种用于反应性对话头部头像的可控3D二元运动先验。STEER将对话行为分解为对目光、头部节奏和情感的明确控制,允许用户引导头像倾听、反应和与对话伙伴互动。由于公共二元语料库中没有这些行为在时间上对齐的注释,我们引入了一个跟踪和注释管道,从自然二元视频中恢复行为伪标签。然后,一个因果流匹配变压器学习基于音频、伙伴运动、情感和提议的行为控制的伙伴感知目标运动。我们通过扩展通用高斯头部头像先验,将STEER进一步嵌入到逼真的头像管道中,通过从跟踪的参数运动到其头像驱动空间的学习映射。这使得无需重新训练底层头像模型就能对高保真高斯头部头像进行可控动画。STEER在运动质量、动力学和多样性方面优于最近的二元运动基线,在伙伴耦合方面保持竞争力,并实现了目光、头部节奏和情感编辑以及交互式实时部署。我们在网页上提供了代码和数据集注释。

英文摘要

Facial movement and expression are central to face-to-face communication, conveying turn-taking, attention, agreement, and engagement alongside speech. While speech-driven facial animation has made strong progress in lip synchronization and audio-conditioned motion generation, most methods treat conversational behavior as an emergent byproduct of audio, or expose only coarse sequence-level affect control. As a result, key non-verbal channels such as gaze contact and aversion, rhythmic head motion, and emotion remain difficult to explicitly control. We present STEER, a controllable 3D dyadic motion prior for reactive conversational head avatars. STEER factorizes conversational behavior into explicit controls for gaze, head rhythm, and emotion, allowing users to steer how an avatar listens, reacts, and engages with a conversation partner. Since temporally aligned annotations for these behaviors are not available in public dyadic corpora, we introduce a tracking and annotation pipeline that recovers behavioral pseudo-labels from in-the-wild dyadic video. A causal flow-matching transformer then learns partner-aware target motion conditioned on audio, partner motion, emotion and the proposed behavioral controls. We further embed STEER in a photorealistic avatar pipeline by extending a Universal Gaussian Head-Avatar Prior with a learned mapping from tracked parametric motion into its avatar-driving space. This enables controllable animation of high-fidelity Gaussian head avatars without re-training the underlying avatar model. STEER outperforms recent dyadic motion baselines on motion quality, dynamics, and diversity, remains competitive on partner coupling, and enables gaze, head-rhythm, and emotion edits together with an interactive live deployment. We make our code and dataset annotations available at our webpage.

CommentsProject page: https://kartik-teotia.github.io/STEER/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑