arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GLID:门控局部内在维度修复面部伪造检测器的盲点

GLID: Gated Local Intrinsic Dimension Repairs the Blind Spots of Face-Forgery Detectors

Guang Yang, Fengchen Liu

arXiv 2607.18770首次发表:更新:

发表机构

University of California, Los Angeles; University of California, Berkeley(加州大学洛杉矶分校; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对微调基础模型面部伪造检测器对训练外生成器家族有盲点的问题,提出GLID检测器,通过估计图像补丁令牌局部内在维度,经置信门进入微调检测器,在基准测试中表现出色,揭示相关经验法则并降低准确率跨种子传播。

AI 中文摘要

微调后的基础模型检测器在面部伪造基准测试中占主导地位,但对训练中未出现的生成器家族视而不见。我们提出了GLID,一种用几何而非数据修复此盲点的检测器。GLID将单个图像的补丁令牌视为来自流形的样本,并在冻结视觉变压器的几个深度估计其局部内在维度(LID)。这个12维、无需训练的信号通过一个置信门进入微调检测器,该门的强度仅在分布内校准。在16轴交叉生成器基准测试中,GLID的平均AUC达到0.805,在重新训练的最先进基线中排名第一,在任何轴上都从未显著落后于最强基线。它将生成轴的AUC提高了+0.084,而将重演轴仅降低了-0.005。两条经验法则解释了该设计。首先,伪造面部在特定家族深度弯曲令牌流形:GAN伪像在最后一层达到峰值,扩散伪像在网络中间达到峰值,并且这种模式在四个主干、三个维度估计器和非面部图像中都存在。其次,微调恰好在训练数据覆盖的地方吸收辅助增益:注入1%的目标家族图像会消除+0.100的增益,因此几何信号在数据不可用的地方至关重要。确定性信号还将准确率的跨种子传播降低了5.5倍。论文还附带了代码、预注册分析门和每张图像的分数。

英文摘要

Fine-tuned foundation-model detectors dominate face-forgery benchmarks, yet they stay blind to generator families absent from training. We present GLID, a detector that repairs this blind spot with geometry instead of data. GLID treats the patch tokens of a single image as a sample from a manifold and estimates their local intrinsic dimension (LID) at several depths of a frozen vision transformer. This 12-dimensional, training-free signal enters a fine-tuned detector through a confidence gate whose strength is calibrated purely in-distribution. On a 16-axis cross-generator benchmark, GLID reaches 0.805 mean AUC, first among retrained state-of-the-art baselines and never significantly behind the strongest of them on any axis. It lifts the generation axes by +0.084 AUC while moving reenactment by only -0.005. Two empirical laws explain the design. First, forged faces bend the token manifold at family-specific depths: GAN artifacts peak at the last layer, diffusion artifacts peak mid-network, and the pattern survives four backbones, three dimension estimators, and non-face imagery. Second, fine-tuning absorbs auxiliary gains exactly where training data covers: injecting 1% target-family images erases a +0.100 gain, so geometric signals matter precisely where data is unavailable. The deterministic signal also cuts the cross-seed spread of accuracy 5.5x. Code, preregistered analysis gates, and per-image scores accompany the paper.

Comments21 pages, 16 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑