发表机构
International Digital Economy Academy(国际数字经济研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有方法在关节化人像网格估计中的局限,GRAPE构建人像参数模型,提出渐进式解剖对齐网络,经多源监督训练,提升了人像网格恢复质量、姿态对齐及颌部与表情解缠能力,还惠及下游任务。
AI 中文摘要
关节化人像网格估计对于3D理解、虚拟化身生成和沉浸式交互至关重要。现有方法主要依赖3D可变形模型(3DMMs),但存在诸多局限。为解决这些问题,我们引入GRAPE。构建了具有明确躯干到头部运动链的人像参数模型(PPM),提出了渐进式解剖对齐(PAA)网络,由预训练人像编码器、渐进式掩码路由器和遵循人像解剖先验的粗细专家组成,并采用多源监督训练。实验表明GRAPE在多方面优于现有方法,还能惠及下游任务。
英文摘要
Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rely on 3D Morphable Models (3DMMs). However, face-centric models suffer from the "floating head" assumption, conflating head pose with global rotation due to the lack of neck kinematics. Conversely, body-centric models lack high-fidelity facial expression capabilities. Furthermore, current methods struggle to disentangle jaw articulation from expression blendshapes, often over-relying on expressions for mouth opening. These limitations make monocular portrait recovery difficult across representation, supervision, and anatomical parameter estimation. To address these limitations, we introduce GRAPE(Graduated Routing for Articulated Portrait mesh Estimation). We build a Portrait Parametric Model (PPM) with an explicit torso-to-head kinematic chain and a canonical injection step to merge FLAME and the SMPL-X torso. We propose a Progressive Anatomical Alignment (PAA) network, which is composed of a pretrained portrait encoder, a Graduated-Mask Router, and coarse-to-fine experts that follow the portrait anatomical prior. We then train this network with multi-source supervision that combines sparse anatomical keypoints, feature distillation, foreground mask constraints, and relative geometry constraints. Experiments show that GRAPE improves portrait mesh recovery quality, pose alignment, and jaw--expression disentanglement over prior methods. We also demonstrate that our method can benefit the downstream tasks of audio-driven talking-head generation and 3D portrait generation.