arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07863cs.CVeess.IV

LHSDet:基于视觉问答的高分辨率AI生成图像检测方法

LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering

Qian Yao, Jun-Jie Huang, Yongjun Wang, Luming Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有AI生成图像检测方法忽略高分辨率细节、难以应对未知生成模型的问题,提出LHSDet,将检测任务转为视觉问答,采用三分支架构,在各类生成模型上实现高检测准确率与稳健性能。

中文摘要 AI 辅助

受扩散模型和自回归模型进展的驱动,AI生成图像的保真度和分辨率现已可与真实图像相媲美。然而,现有的AI生成图像检测方法常对图像进行下采样,不可避免地忽略了高分辨率AI生成图像中的关键低级纹理细节,从而限制了其检测性能。此外,未知生成模型的不断涌现使得大规模预训练数据集难以获取。为应对这些挑战,我们提出了一种新型高分辨率AI生成图像检测器,命名为LHSDet。具体而言,我们将AI生成图像检测任务形式化为视觉问答问题,利用微调后的视觉-语言框架充分挖掘视觉与文本模态之间的互补信息。考虑到现有视觉-语言模型的默认视觉编码器并非为AI生成图像检测量身定制,我们重新设计了视觉编码器,以更好地捕捉AI生成图像中固有的低级和高级伪影。此外,我们引入了语义级文本分支以实现多模态特征融合与检测。因此,LHSDet采用三分支架构提取互补的多模态特征:用于聚合非重叠补丁以获取局部纹理线索的低级视觉分支、基于SigLIP2用于全局感知特征提取的高级视觉分支,以及使用BLIP-2生成描述的语义级文本分支。大量实验结果表明,LHSDet在包括扩散模型和自回归模型在内的各类生成模型上均实现了高检测准确率和稳健性能。

英文摘要

Driven by advances in diffusion models and autoregressive models, the fidelity and resolution of AI-generated images now rival those of real images. However, existing AI-generated image detection methods often downsample the images, inevitably overlooking critical low-level texture details in high-resolution AI-generated images, therefore limiting their detection performance. In addition, the ceaseless emergence of unknown generative models makes large-scale pre-training datasets inaccessible. To address these challenges, we propose a novel high-resolution AI-generated image detector, termed LHSDet. Specifically, we formulate the AI-generated image detection task as a Visual Question Answering problem, leveraging a fine-tuned vision-language framework to fully exploit the complementary information between visual and textual modalities. Recognizing that the default visual encoder of existing vision-language models is not tailored for AI-generated image detection, we redesign a visual encoder to better capture both the low-level and high-level artifacts inherent in AI-generated images. Furthermore, we incorporate a semantic-level textual branch to enable multi-modal feature fusion and detection. Consequently, LHSDet employs a triple-branch architecture to extract complementary multi-modal features: a low-level visual branch that aggregates non-overlapping patches for local texture cues, a high-level visual branch based on SigLIP2 for global perception feature extraction, and a semantic-level textual branch that generates captions using BLIP-2. Extensive experimental results demonstrate that LHSDet achieves high detection accuracy and robust performance across diverse generative models, including both diffusion and autoregressive models.

发表机构

  • College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院)
  • Academy of Military Science(军事科学院)

机构由 AI 辅助整理,请以论文原文为准。

↑