发表机构
People’s Daily Online; Nanyang Technological University; Tongji University; Nanjing University of Science and Technology; National University of Singapore; University of Science and Technology of China(人民网; 南洋理工大学; 同济大学; 南京理工大学; 新加坡国立大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有生成模型图像检测基准与现实差距问题,提出统一基准DailyBench,含FakeBench和ManipulationBench两个子集,通过实验揭示当前检测器鲁棒性差,凸显其可用于开发更强大的人工智能生成图像检测方法。
AI 中文摘要
生成模型的最新进展使人工智能生成图像检测从识别易于区分的完全合成图像,转向识别由现代生成和处理管道生成的高度逼真的内容。然而,现有的检测基准通常基于过时的生成模型构建,主要强调全图像合成,导致基准数据与现实世界生成和编辑场景中遇到的图像之间的差距越来越大。为弥合这一差距,我们引入了DailyBench,这是一个高质量的统一基准,用于评估人工智能生成图像检测器能否在现代全图像合成和对象级处理中进行泛化。DailyBench包含两个互补子集:FakeBench,包括由近期开源和商业生成模型合成的高质量图像;ManipulationBench,引入了使用先进图像条件模型应用于真实图像的具有挑战性的对象级编辑。这种设计使DailyBench成为研究生成器级泛化和在细微局部编辑下的操作感知检测的现实测试平台。在DailyBench上的实验揭示了当前检测器存在巨大的鲁棒性差距:在GenImage上报告平衡准确率为91-96%的方法,在FakeBench上降至60-76%,在ManipulationBench上降至54-66%。这些结果表明,现有检测器对现实合成和处理仍缺乏泛化能力,凸显了DailyBench作为开发强大且具有操作感知能力的人工智能生成图像检测方法的严格测试平台的地位。该项目可通过此https链接获取。
英文摘要
Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identifying highly realistic content generated by both modern generation and manipulation pipelines. However, existing detection benchmarks are often built with outdated generative models and primarily emphasize full-image synthesis, creating a growing mismatch between benchmark data and the images encountered in real-world generation and editing scenarios. To bridge this gap, we introduce DailyBench, a high-quality unified benchmark for evaluating whether AI-generated image detectors can generalize across both modern full-image synthesis and object-level manipulation. DailyBench contains two complementary subsets: FakeBench, which includes high-quality images synthesized by recent open-source and commercial generative models, and ManipulationBench, which introduces challenging object-level edits applied to real images using advanced image-conditional models. This design makes DailyBench a realistic testbed for studying both generator-level generalization and manipulation-aware detection under subtle local edits. Experiments on DailyBench reveal substantial robustness gaps in current detectors: methods reporting 91-96% balanced accuracy on GenImage drop to 52-79% on FakeBench and 43-67% on ManipulationBench. These results show that existing detectors remain poorly generalized to realistic synthesis and manipulation, highlighting DailyBench as a rigorous testbed for developing robust and manipulation-aware AI-generated image detection methods. The project is available at https://dailybench.github.io/
CommentsSome errors have been fixed; please refer to the latest submitted version