可重光照的3D头像重建:语义自适应的运动-光照响应
Relightable 3D Avatar Reconstruction with Semantic-Adaptive Motion-Illumination Responses
浏览论文内容
中文总结 AI 辅助
针对现有高斯头像统一建模忽略面部语义差异的问题,提出SAMIRA框架,通过语义自适应运动与光照响应模块,提升细粒度表情重建和重光照逼真度。
中文摘要 AI 辅助
从单目视频重建富有表现力且可重光照的3D头部头像仍然是计算机视觉中的一项挑战,因为它需要对非刚性面部运动和光照相关外观进行精确建模。现有的高斯头像方法通常依赖于全局耦合表示,其中高斯原语共享统一的运动或光照响应模型。这种统一建模忽略了不同面部语义区域的不同运动模式和材质/反射属性,从而限制了细粒度动画的准确性并降低了重光照的逼真度。为解决这一局限,我们提出了SAMIRA,一个用于语义自适应运动-光照响应建模的3D高斯头像框架。对于运动响应建模,语义自适应运动响应模块将当前到参考的网格位移光栅化到拓扑一致的UV空间中,并利用面部语义将位移特征路由到语义特定的调制器中,预测超出粗略网格绑定的局部化高斯几何残差。对于光照响应建模,语义自适应光照响应模块为每个面部区域学习紧凑的漫反射和镜面反射响应因子,使不同区域的高斯能够适应新颖环境光照的照明响应。这些响应因子被纳入延迟的基于物理的着色中,提供了语义相关光照效应的轻量级近似。在自重演、交叉重演和重光照上的大量实验表明,与现有方法相比,SAMIRA提高了细粒度表情重建和重光照逼真度。
英文摘要
Reconstructing expressive and relightable 3D head avatars from monocular videos remains challenging in computer vision, as it requires accurate modeling of both non-rigid facial motion and illumination-dependent appearance. Existing Gaussian avatar methods commonly rely on globally coupled representations, in which Gaussian primitives share a unified motion or illumination response model. Such uniform modeling neglects the distinct motion patterns and material/reflectance properties of different facial semantic regions, thereby limiting fine-grained animation accuracy and reducing relighting plausibility. To address this limitation, we propose SAMIRA, a 3D Gaussian avatar framework for semantic-adaptive motion-illumination response modeling. For motion response modeling, the Semantic-Adaptive Motion Response module rasterizes current-to-reference mesh displacements into a topology-consistent UV space and leverages facial semantics to route displacement features through semantic-specific modulators, predicting localized Gaussian geometric residuals beyond coarse mesh binding. For illumination response modeling, the Semantic-Adaptive Illumination Response module learns compact diffuse and specular response factors for each facial region, allowing Gaussians in different regions to adapt their illumination responses to novel environment lighting. These response factors are incorporated into deferred physically based shading, providing a lightweight approximation of semantic-dependent illumination effects. Extensive experiments on self-reenactment, cross-reenactment, and relighting demonstrate that SAMIRA improves both fine-grained expression reconstruction and relighting realism over existing methods.
发表机构
- State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统全国重点实验室)
- School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
- ZKTeco Co., Ltd.(熵基科技股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。