AI 中文总结
REALIS是一个包含143万张图像、由42个生成器创建的数据集,用于评估AI图像检测器在内容多样性和后处理下的鲁棒性,并引入专家子集和变换协议,最佳检测器ROC-AUC从0.550提升至0.752。
AI 中文摘要
AI生成的图像检测器通常在真实图像与合成图像在内容、质量或生成伪影方面存在差异的基准上进行评估,这使得模型能够依赖数据集特定的线索,并在面对不熟悉的生成器或经过处理的图像时失效。现有数据集在跨多样视觉内容联合评估这些挑战方面提供的支持有限。我们引入了REALIS,这是一个包含143万张真实和合成图像的数据集,由42个现代文本到图像模型生成,包括最新的专有系统如Nano Banana 2。REALIS结合了源自真实图像的提示、质量过滤和分层采样,以减少类别特定的捷径,同时保持内容多样性。我们进一步引入了REALIS-Expert,一个用于高质量合成图像的压力测试子集,其中真实和生成的样本在语义和视觉特征上被紧密匹配。我们还提出了一个鲁棒性协议,涵盖五个严重级别的35种变换,以分析检测器在图像处理下的行为。基于REALIS,我们的基准在生成器和后处理偏移下评估了预训练检测器、微调模型和零样本视觉语言模型。在最难处理的划分上,最佳预训练传统检测器实现了0.550的ROC-AUC,而最佳REALIS训练检测器则为0.752。REALIS提供了一个统一框架,用于在更贴近实际使用的条件下测量和提高AI图像检测器的可靠性。
英文摘要
AI-generated image detectors are often evaluated on benchmarks where real and synthetic images differ in content, quality, or generation artifacts, allowing models to rely on dataset-specific cues and fail on unfamiliar generators or processed images. Existing datasets provide limited support for evaluating these challenges jointly across diverse visual content. We introduce REALIS, a dataset of 1.43 million real and synthetic images generated by 42 modern text-to-image models, including the latest proprietary systems such as Nano Banana 2. REALIS combines prompts derived from real images, quality filtering, and stratified sampling to reduce class-specific shortcuts while preserving content diversity. We further introduce REALIS-Expert, a stress-test subset for high-quality synthetic images, where real and generated samples are selected with closely matched semantic and visual characteristics. We also propose a robustness protocol covering 35 transformations at five severity levels to analyze detector behavior under image processing. Based on REALIS, our benchmark evaluates pretrained detectors, fine-tuned models, and zero-shot vision-language models under generator and post-processing shifts. On the hardest processed split, the best pretrained conventional detector achieves 0.550 ROC-AUC, compared with 0.752 for the best REALIS-trained detector. REALIS provides a unified framework for measuring and improving the reliability of AI-image detectors under conditions that better reflect real-world use.