发表机构
The Hong Kong Polytechnic University; Macao Polytechnic University(香港理工大学; 澳门理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DisFace3DNet通过3D组件解耦技术,在SCUT-FBP5500数据集上实现与人类评分皮尔逊相关系数0.8904±0.0063的面部吸引力预测,可解释各组件贡献并连接整体预测与定量分析。
AI 中文摘要
面部吸引力预测通常仅给出一个整体评分,而将形状、外观和观看条件的作用隐含其中。我们提出DisFace3DNet,该模型利用3D组件解耦技术,在无需人工标注组件目标的情况下,借助辅助弱语义监督从整体评分中学习七个组件的参考评分。指定的3D表示和图像线索共同输入到针对身份、皮肤、头发、光照、背景、表情和姿态的联合学习路径中。随后通过约束拟合将五个静态评分和两个带符号的动态评分组合为整体评分,从而揭示每个组件的数值贡献,并支持跨图像的组件特定比较。在SCUT-FBP5500数据集上,DisFace3DNet与平均人类评分的皮尔逊相关系数达到0.8904±0.0063(五折交叉验证的均值±标准差);其组件项可将每个留出预测重构至数值精度。皮肤、头发和面部形状是组件级预测变化最大的因素。人类评估支持面部形状、皮肤和头发的评分方向;表情的一致性较弱。因此,DisFace3DNet将整体预测与构成每个估计的面部及上下文线索的定量分析相连接。
英文摘要
Facial attractiveness prediction usually assigns one overall rating, leaving the roles of shape, appearance, and viewing conditions implicit. We propose DisFace3DNet, which uses 3D component disentanglement to learn seven component reference scores from overall ratings with auxiliary weak semantic supervision, without human-labeled component targets. Designated 3D representations and image cues feed jointly learned routes for identity, skin, hair, light, background, expression, and pose. A constrained fit then combines five static and two signed dynamic scores into the overall rating, exposing each component's numerical contribution and supporting component-specific comparisons across images. On SCUT-FBP5500, DisFace3DNet achieves a Pearson correlation of $0.8904\pm0.0063$ (mean $\pm$ standard deviation across five folds) with average human ratings; its component terms reconstruct every held-out prediction to numerical precision. Skin, hair, and facial shape account for the largest component-wise prediction variation. Human evaluation supports the score directions for facial shape, skin, and hair; expression agreement is weaker. DisFace3DNet thus connects overall prediction to quantitative analysis of the facial and contextual cues entering each estimate.
CommentsIncludes supplemental materials