arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11134cs.CV

LAION-Mobile:在百万智能手机照片上评估深度伪造检测器

LAION-Mobile: Evaluating Deepfake Detectors On One Million Smartphone Photos

Achim von Stryk, Janis Keuper

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出LAION-Mobile数据集,含百万智能手机照片,评估十二种深度伪造检测器,发现现代AI内容上性能低下且真实照片误报率高,揭示现有检测器在计算摄影时代的局限性。

中文摘要 AI 辅助

大多数深度伪造检测器在其参考基准上报告接近完美的AUC分数。然而,最近一篇ICML立场论文认为,这些评估共同忽略了现代智能手机摄影的影响:广泛使用的设备端神经图像信号处理流水线(如多传感器融合或噪声和运动模糊抑制)日益将成像范式从简单的镜头投影转向计算摄影。因此,设备实际上生成而非记录照片。这增加了深度伪造检测器可能将普通手机照片标记为伪造的风险。由于缺乏包含现代智能手机图像的大规模数据集,这一假设迄今为止仅在小型概念验证研究中得到测试。本文旨在弥补这一差距。我们引入了LAION-Mobile,这是一个开放数据集,包含约100万张带有EXIF元数据的智能手机图像,这些图像从re-LAION-5B中提炼而来。我们使用十二种最先进的深度伪造检测器及其原始论文检查点,在该数据池的9,115张图像评估样本(DIRE为738张)上进行评估,报告了三个关键发现:(i)在现代AI内容上,没有检测器的AUC超过0.624,十二种中有五种低于随机水平。(ii)真实照片的误报率是阈值校准的产物:在传统GAN数据上拟合的阈值使几种检测器看起来可部署(误报率低于11%),然而一旦在现代内容上重新拟合相同标准,这些检测器会将17%至91%的真实照片标记为伪造。(iii)因此,没有检测器能同时在现代AI内容上超过随机水平并保持可部署的真实照片误报率。该语料库反映了网络集合的设备构成,探测了第一代神经ISP(2018-2020);当前旗舰机型基本缺失,留下现代ISP机制作为开放空白。

英文摘要

Most Deepfake detectors report near-perfect AUC scores on their reference benchmarks. However, a recent ICML position paper argues that these evaluations collectively neglect the impact of modern smartphone photography: the widely used on-device neural image-signal processing pipelines (like multi-sensor fusion or noise and motion-blur suppression) increasingly shift the imaging paradigm from simple lens projections towards computational photography. Hence, devices actually generate, rather than record photos. This increases the risk that deepfake detectors may flag ordinary phone photos as fake. Due to the lack of large-scale datasets containing images from modern smartphones, this hypothesis has so far only been tested in small proof-of-concept studies. The aim of this paper is to close this gap. We introduce LAION-Mobile, an open dataset containing about 1 million smartphone images with EXIF metadata distilled from re-LAION-5B. Evaluating twelve state-of-the-art deepfake detectors with their original paper checkpoints on a 9,115-image evaluation sample of this pool (DIRE on 738), we report three key findings: (i) On modern AI content no detector exceeds AUC 0.624, and five of twelve fall below chance. (ii) Real-photo false-alarm rates are an artefact of threshold calibration: thresholds fitted on legacy GAN data make several detectors look deployable (less than 11 percent FPR), yet the same detectors flag 17-91 percent of real photos once the identical criterion is refit on modern content. (iii) Consequently, no detector both beats chance on modern AI content and keeps a deployable real-photo false-alarm rate. Mirroring the device mix of web collections, the corpus probes the first neural-ISP generation (2018-2020); current flagships are essentially absent, leaving the modern-ISP regime as the open gap.

发表机构

  • Stralsund University(施特拉尔松德大学)
  • IMLA, Offenburg University(奥芬堡大学IMLA)

机构由 AI 辅助整理,请以论文原文为准。

↑