arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23984cs.CV

基于单张肖像重建的3D高斯头部的源人脸真实性检测:基准与专用检测器

Source-Face Authenticity Detection for 3D Gaussian Heads Reconstructed from a Single Portrait: A Benchmark and Dedicated Detector

  • Shanghai Jiaotong University(上海交通大学)
  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

Yujie Gao, Zijian Yu, Yan Hong, Jun Lan, Jianfu Zhang

AI总结:

针对单张肖像重建的3D高斯头部的源人脸真实性检测问题,构建首个大规模基准并发现现有检测器的局限,提出两阶段训练的专用检测器,在所有指标上实现最高准确率。

AI中文摘要:

近期,单图像3D高斯头部重建技术取得进展,可从单张肖像生成高度逼真且可自由渲染的数字头部。但重建与渲染会削弱源肖像中的伪造痕迹,导致生成的3D人脸难以判断其底层人脸是真实还是伪造,对身份认证与人脸隐私构成风险。为研究该问题,我们构建了首个针对此任务的大规模基准,从多源收集真实肖像与伪造肖像,并在该基准上评估现有代表性检测器,发现它们缺乏保留细粒度信息、维持渲染视图间特征一致性的明确机制。为直接解决这两个局限,我们提出采用两阶段策略训练的检测器:阶段I中,掩码自编码鼓励视觉骨干保留局部重建所需的细粒度外观信息,多视图对比学习则强制同一头部的渲染视图间特征一致;因不同深度的CLS token呈现互补的空间注意力模式,阶段II冻结适配后的骨干,拼接低、中、高层CLS token用于分类。实验表明,在所有评估检测器中,我们的方法在报告的所有指标上均达到最高准确率,排名第一。

英文摘要:

Recent advances in single-image 3D Gaussian head reconstruction have enabled highly realistic and freely renderable digital heads from a single portrait. However, reconstruction and rendering can weaken the forgery traces in the source portrait, making the resulting 3D face difficult to classify whether its underlying face is real or fake, and thereby posing risks to identity authentication and face privacy. To study this problem, we introduce the first large-scale benchmark for this task by collecting real portraits and fake portraits from multiple sources and evaluate representative existing detectors on this benchmark, revealing their lack of explicit mechanisms for retaining fine-grained information and maintaining feature consistency across rendered views. To directly address these two limitations, we propose a detector trained with a two-stage strategy. In Stage I, masked autoencoding encourages the visual backbone to retain the fine-grained appearance information required for local reconstruction, while multi-view contrastive learning enforces feature consistency across rendered views of the same head. Since CLS tokens at different depths exhibit complementary spatial attention patterns, Stage II freezes the adapted backbone and concatenates low-, middle-, and high-level CLS tokens for classification. Experiments show that our method achieves the highest accuracy and ranks first across all reported metrics among the evaluated detectors.

↑