arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GHARP:基于大规模重建先验的实时高斯头部动画

GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior

Ali Benlalah, Sepehr Johari, Patricia Vitoria, Armin Kappeler, Artem Sevastopolsky, Alexander Jung, Gabriele Fanelli, Kevin Mader, Manuel Breitenstein, Claudia Plüss, Jan Rüegg, Simon Biland, Thomas Etterlin, Dmitry Kostiaev, Mathias Deschler, Brian Amberg, Sebastian Martin

arXiv 2610.10945首次发表:更新:

发表机构

Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GHARP是一种实时高斯头部动画方法,通过解耦身份与动画阶段、引入身体对齐网络,在Ava-256基准上实现最优质量,且在A100 GPU上速度快13倍、高斯数减8倍。

AI 中文摘要

我们提出了GHARP(Real-time Gaussian Head Animation from Large-scale Reconstruction Prior),这是一种从目标对象的少量输入图像和驱动表情信号实时生成3D人头部动画的方法。我们将该问题解耦为两个阶段:身份阶段离线构建对象的几何与外观表示,以及动画阶段在运行时基于该表示预测依赖于表情的残差。这种分离在保真度、质量和运行速度之间实现了良好的权衡:身份阶段可以是计算密集型的,而动画阶段则运行一个针对移动设备优化的轻量级网络。我们的方法在预训练重建模型的语义结构化潜在空间中执行动画,其中表情变化保持空间上的一致性,使残差预测变得高效。该重建先验提供了一致的空间布局,允许将多个输入视图融合为紧凑、固定大小的规范高斯表示。尽管这种两阶段设计改善了运行速度与质量的权衡,但它仍然继承了所有表情驱动化身方法共有的问题:表情编码仅描述面部,而忽略了身体姿态和衣物位置,导致这些区域在输入中未被充分指定。动画网络面临不适定映射,只能在冲突的身体外观上取平均,从而产生模糊和时间闪烁。我们通过一个身体对齐网络解决了这个问题,该网络学习将目标图像中的人物身体与输入参考图像对齐,消除了训练信号中的歧义。我们的方法在Ava-256基准上达到了最先进的质量,同时在A100 GPU上运行速度快达13倍,且高斯数量减少了8倍。

英文摘要

We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.

CommentsAccepted to ACCV 2026. 35 pages (14 main + references + 15 pages supplementary), 14 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑