发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GHARP是一种实时高斯头部动画方法,通过解耦身份与动画阶段、引入身体对齐网络,在Ava-256基准上实现最优质量,且在A100 GPU上速度快13倍、高斯数减8倍。
AI 中文摘要
我们提出了GHARP(Real-time Gaussian Head Animation from Large-scale Reconstruction Prior),这是一种从目标对象的少量输入图像和驱动表情信号实时生成3D人头部动画的方法。我们将该问题解耦为两个阶段:身份阶段离线构建对象的几何与外观表示,以及动画阶段在运行时基于该表示预测依赖于表情的残差。这种分离在保真度、质量和运行速度之间实现了良好的权衡:身份阶段可以是计算密集型的,而动画阶段则运行一个针对移动设备优化的轻量级网络。我们的方法在预训练重建模型的语义结构化潜在空间中执行动画,其中表情变化保持空间上的一致性,使残差预测变得高效。该重建先验提供了一致的空间布局,允许将多个输入视图融合为紧凑、固定大小的规范高斯表示。尽管这种两阶段设计改善了运行速度与质量的权衡,但它仍然继承了所有表情驱动化身方法共有的问题:表情编码仅描述面部,而忽略了身体姿态和衣物位置,导致这些区域在输入中未被充分指定。动画网络面临不适定映射,只能在冲突的身体外观上取平均,从而产生模糊和时间闪烁。我们通过一个身体对齐网络解决了这个问题,该网络学习将目标图像中的人物身体与输入参考图像对齐,消除了训练信号中的歧义。我们的方法在Ava-256基准上达到了最先进的质量,同时在A100 GPU上运行速度快达13倍,且高斯数量减少了8倍。
英文摘要
We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.
CommentsAccepted to ACCV 2026. 35 pages (14 main + references + 15 pages supplementary), 14 figures, 17 tables