发表机构
École de Technologie Supérieure; Ubisoft La Forge(高等技术学院; 育碧拉福奇工作室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FaceSnap是端到端框架,通过两阶段方法实现单目灯光舞台相机的实时个性化人脸表演捕捉,几何精度媲美逐帧多视图优化,优于前馈方法,还引入了4D人脸重建基准Multi4D。
AI 中文摘要
灯光舞台人脸捕捉可生成制作级数字人,但资源与人力密集,多相机设置、数小时计算及海量数据存储形成瓶颈,阻碍迭代工作流。本文提出FaceSnap,一种端到端框架,通过两阶段方法简化捕捉:首先,从动作范围序列进行一次性多视图优化,构建编码几何与表情依赖外观的个性化模型;该模型随后可从单目灯光舞台相机实现高保真实时人脸表演捕捉,无需进一步多视图捕捉。FaceSnap以83 fps联合估计几何与4K动态纹理,4K纹理由新型个性化残差上采样器生成,可恢复主体特定高频细节,通用上采样器无法捕获。FaceSnap几何精度可与逐帧全多视图优化媲美,且优于在制作级3D数据上训练的前馈方法,所有结果均来自单相机视角。最后,本文引入Multi4D,一个用于评估灯光舞台环境中4D人脸重建方法的公开基准,支持跨方法的拓扑不变几何比较。
英文摘要
Lightstage facial capture produces production-quality digital humans, but it is resource and labor-intensive. Multi-camera setups, hours of computation, and massive data storage create bottlenecks that hinder iterative workflows. This paper introduces FaceSnap, an end-to-end framework that streamlines capture via a two-stage approach. First, a one-time multi-view optimization from a range-of-motion sequence builds a personalized model encoding both geometry and expression-dependent appearance. This model then enables high-fidelity real-time facial performance capture from a single monocular lightstage camera, with no further multi-view capture required. FaceSnap jointly estimates geometry and dynamic 4K texture at 83 fps. The 4K texture is produced by a novel personalized residual upscaler that recovers subject-specific high-frequency detail, which generic upscalers fail to capture. FaceSnap achieves geometric accuracy competitive with full per-frame multi-view optimization while outperforming feed-forward methods trained on production-quality 3D data, all from a single camera view. Finally, we introduce Multi4D, a public benchmark for evaluating 4D facial reconstruction methods in lightstage environments, enabling topology-invariant geometric comparison across methods.