arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面部追踪:合成面部生成器的开放集归因与渐进式发现

Face-Trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators

Alessia Infantino, Claudio Schiavella, Irene Amerini

arXiv 2607.07545首次发表:更新:

发表机构

Department of Computer, Control and Management Engineering (DIAG), Sapienza University of Rome(罗马第一大学计算机、控制与管理工程系(DIAG))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对合成面部图像逼真带来的多媒体取证挑战,提出结合已知生成器分类、异常检测拒绝及未知生成器发现的开放集归因管道,实验验证其在封闭集和开放集设置下的有效性及增量设置下的渐进式发现能力,跨数据集实验显示可超越原数据集分布。

AI 中文摘要

生成式人工智能的进展使合成面部图像更逼真,给多媒体取证带来挑战。现有方法多在封闭集设置下处理合成面部归因,而实际新生成器不断出现。本文提出开放集合成面部源归因管道,结合已知生成器分类、基于能量的异常检测拒绝和未知生成器发现。在已知生成器上训练分类器,对拒绝样本聚类。实验表明该方法在封闭集归因准确率达96.73%,开放集下基于能量的拒绝平衡准确率达71.25%,拒绝样本聚类效果良好,增量设置下发现的生成器空间逐步扩展且最终纯度达99.23%,跨数据集实验表明管道可超越原始数据集分布,但后处理仍具挑战。

英文摘要

Recent advances in generative Artificial Intelligence have made synthetic face images increasingly realistic, creating new challenges for multimedia forensics. Source attribution methods should identify the generator of an image when the source is known, but also handle samples produced by unseen models. Most existing approaches, however, address synthetic face attribution in a closed-set setting, assuming that test samples can only originate from generators observed during training. This assumption does not hold in real-world scenarios, where new generators continuously appear and detecting an image as unknown is not sufficient, since rejected samples should also be organized according to their underlying sources. We introduce Face-Trace, a pipeline for open-set synthetic face source attribution that combines known generator classification, energy-based rejection, and unknown generator discovery. A classifier trained on frozen I-JEPA embeddings attributes known generators, while rejected samples are represented by combining projected I-JEPA features with complementary forensic traces and grouped to identify coherent sets of samples produced by unknown generators. We also extend the discovery stage to an incremental scenario, where rejected samples arrive over time. Experiments on the WILD dataset show 96.73% closed-set attribution accuracy, while rejection reaches 71.25% balanced accuracy and rejected samples are clustered into meaningful unknown-generator groups, with an Adjusted Rand Index of 0.81, a Normalized Mutual Information of 0.90, and an overall purity of 87.74%. In the incremental setting, the discovered generator space is progressively extended while maintaining a final purity of 99.23%, and cross-dataset experiments suggest that the pipeline can operate beyond the original data distribution.

CommentsPreprint. 13 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑