arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AvatarDynamizer:通过生成式动态纹理从静态化身到动态人类化身

AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures

Guoxing Sun, Heming Zhu, Linjie Lyu, Pascal Fua, Christian Theobalt, Marc Habermann

arXiv 2608.19900首次发表:更新:

发表机构

Max Planck Institute for Informatics; EPFL; VIA Research Center(马克斯·普朗克信息学研究所; 洛桑联邦理工学院; VIA研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AvatarDynamizer是一种生成式方法,可将现成静态3D化身转为可控多视角一致4D化身,通过纹理空间动态嵌入实现真实表面动态,在视觉保真度上优于通用方法。

AI 中文摘要

对于全身化身而言,建模表面动态对于克服恐怖谷效应并实现感知真实感至关重要。与人物无关的方法可从单目图像、视频或文本提示中恢复静态3D化身,但它们的骨骼驱动动画缺乏衣物褶皱等真实表面动态;而与人物相关的方法虽能实现高质量渲染和真实动态,但需为每个人进行昂贵的多视角捕捉。近期的通用动态化身方法难以嵌入表面动态,导致多视角一致性或动态表现力受限。为此,我们提出AvatarDynamizer,这是一种将现成静态3D化身转换为可控、真实且多视角一致的4D化身的生成式方法。我们引入一种新颖的纹理空间表面动态嵌入,将化身动态建模表述为条件纹理生成;我们的编码器-解码器表示将姿态依赖动态嵌入动态纹理图,使其能与预训练视频扩散模型兼容,同时将其解码为3D高斯以实现多视角一致渲染。由于现有数据集在规模、序列长度或动作多样性上存在局限,我们收集了一个大规模多视角数据集,包含覆盖多样骨骼动作和表面动态的长序列。实验表明,我们的方法能有效为静态化身赋予真实表面动态,且在视觉保真度上优于现有通用方法,尤其在动态训练数据有限的情况下表现更优。

英文摘要

For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles. In contrast, person-specific methods achieve high-quality rendering and realistic dynamics, but require expensive multi-view captures for each individual. Recent generalizable dynamic avatar methods struggle to embed surface dynamics, leading to either limited multi-view consistency or dynamic expressiveness. To this end, we propose AvatarDynamizer, a generative method that transforms an off-the-shelf static 3D avatar into a controllable, realistic, and multi-view-consistent 4D avatar. We introduce a novel texture-space surface-dynamics embedding and formulate avatar dynamics modeling as conditional texture generation. Our encoder--decoder representation embeds pose-dependent dynamics into dynamic texture maps, enabling compatibility with pre-trained video diffusion models while decoding them into 3D Gaussians for multi-view consistent rendering. Since existing datasets are limited in scale, sequence length, or motion diversity, we collect a large-scale multi-view dataset with long sequences covering diverse skeletal motions and surface dynamics. Experiments show that our method effectively animates static avatars with faithful surface dynamics and outperforms competing generalizable methods in visual fidelity, especially under limited dynamic training data.

CommentsProject page: https://vcai.mpi-inf.mpg.de/projects/AvatarDynamizer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑