arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20922cs.CV

FA-LAM:用于一次性4D可动画化高斯头部的焦点感知大型头像模型

FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

Yingdong Hu, Yisheng He, Yiming Jiang, Zehong Lin, Steven Hoi, Jun Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

提出FA-LAM模型用于一次性创建可动画化高斯头部及实现3D、4D全头部恢复。通过分析注意力机制等确定影响因素,采用对称语义注意力正则化等方法解决问题,增强模型支持多视图和流式4D重建,提升了头部重建质量。

中文摘要 AI 辅助

我们提出了FA-LAM,一种用于一次性创建可动画化高斯头部的焦点感知大型头像模型,同时还能实现静态3D和动态4D全头部恢复。我们方法的核心在于深入分析了注意力机制以及现有最先进方法采用的纠缠重建和动画训练管道。分析确定了影响3D全头部生成质量的两个主要因素:(1)不正确和有噪声的注意力激活,(2)重建和动画任务之间的冲突。为解决第一个问题,我们引入了一种对称和语义注意力正则化策略,利用人头的固有语义和结构对称性。为解开重建和动画的目标,我们开发了一种新颖的双阶段训练管道,将模型的大视角幻觉和动画能力分离到不同模块。此外,我们通过具有定制可见性感知令牌融合的核心自回归修改,增强模型以高效且内存友好的方式支持多视图和流式4D重建。这些创新使FA-LAM能够以卓越质量重建可动画化高斯全头部,特别是在精细面部区域和大视角方面。

英文摘要

We propose FA-LAM, a Focus-Aware Large Avatar Model for one-shot animatable Gaussian head creation, while simultaneously enabling static 3D and dynamic 4D full-head recovery. The core of our method lies in a thorough analysis of the attention mechanisms and the entangled reconstruction and animation training pipeline adopted by prior state-of-the-art approaches. Our analysis identifies two main factors that compromise the quality of 3D full-head generation: (1) incorrect and noisy attention activations, and (2) conflicts between the tasks of reconstruction and animation. To address the first issue, we introduce a symmetric and semantic attention regularization strategy that leverages the inherent semantics and structural symmetry of human heads. To disentangle the objectives of reconstruction and animation, we develop a novel dual-phase training pipeline that separates the model's capabilities for large-view hallucination and animation into distinct modules. Moreover, we enhance our model to support multi-view and streaming 4D reconstruction in an efficient and memory-friendly manner through a core autoregressive modification with tailored visibility-aware token fusion. Collectively, these innovations enable FA-LAM to reconstruct animatable Gaussian full heads with superior quality, particularly in fine facial regions and large viewing angles.

发表机构

  • HKUST(香港科技大学)
  • Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)
  • Beihang University(北京航空航天大学)
  • Lingnan University(岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

↑