发表机构
University of Technology Sydney; Shandong University of Science and Technology(悉尼科技大学; 山东科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对单图像生成可动画头部头像的挑战,提出双轴解耦的SpiD框架,通过内化逐帧驱动、分解面部高斯分支,实现了性能优且速度快的实时高斯头像生成。
AI 中文摘要
从单张图像创建逼真的可动画头部头像仍是数字人合成领域的基础挑战。尽管近期的3D高斯溅射(3D Gaussian Splatting)方法已取得令人满意的结果,但它们依赖外部跟踪管道,其延迟未被纳入推理测量;此外,这些方法采用统一表示,会混淆几何上不同的面部区域,限制了表现力和渲染保真度。我们提出SpiD(Split and Drive),这是一种基于双解耦轴的单图像高斯头像框架。计算轴将逐帧驱动内化,消除了推理时对外部跟踪的依赖;特征轴将头像分解为三个专门的高斯分支,每个分支建模一个几何上不同的面部域。大量实验表明,与最先进的方法相比,该方法始终表现出强劲性能,且在包含完整驱动管道的单GPU上实现了所有对比方法中最快的推理速度。
英文摘要
Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian Splatting methods have achieved promising results, they rely on external tracking pipelines whose latency is excluded from inference measurements. Furthermore, they adopt unified representations that entangle geometrically distinct facial regions, limiting both expressiveness and rendering fidelity. We propose SpiD (Split and Drive), a single-image Gaussian head avatar framework built on two disentanglement axes. The compute axis internalizes per-frame driving, eliminating external tracking dependency at inference. The feature axis decomposes the avatar into three specialized Gaussian branches, each modeling a geometrically distinct facial domain. Extensive experiments demonstrate consistently strong performance against state-of-the-art methods while achieving the fastest inference speed among all compared methods on a single GPU with the complete driving pipeline included.