arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15238cs.CV

UC-VLM:基于一致性驱动学习的视觉语言大模型生成图像检测方法

UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models

Lei Tan, Shuwei Li, Mohan Kankanhalli, Robby T. Tan

首次发表
浏览论文内容

中文总结 AI 辅助

UC-VLM是仅依赖二元监督的多阶段框架,通过复用真实性标签实现视觉适配与文本生成,在GenImage、Chameleon等数据集上显著提升了AI生成图像检测的准确率。

中文摘要 AI 辅助

视觉语言大模型(Vision-Language Large Models, VLLMs)因能同时输出预测结果与自然语言文本,在AI生成图像(AI-Generated Image, AIGI)检测领域具有应用潜力。然而,现有多数基于VLLM的检测器主要对语言侧进行微调,对底层视觉取证线索关注有限,且往往依赖人工设计的提示词或人工标注的理由,限制了其泛化能力。本文提出UC-VLM,一种仅依赖二元监督的AIGI检测统一多阶段框架。UC-VLM首先自动识别有效的指令变体,随后在多阶段训练框架内复用相同的二元标签:(i)视觉判别目标,用于增强对非语义取证线索的敏感性;(ii)标签条件生成目标,利用二元标签监督文本输出。该设计将弱二元监督转化为视觉通路与语言输出的共享监督信号。本文的核心创新在于构建统一的多阶段二元监督框架,该框架一致复用相同的真实性标签以实现视觉适配与标签条件文本生成,同时利用自动优化的指令降低提示词敏感性,无需人工编写的理由或手工设计的提示。实验表明,UC-VLM在GenImage数据集上的平均准确率达96.1%,超出现有最强结果4.6%;在Chameleon数据集上,针对ProGAN训练生成的图像准确率为69.6%,针对SDV1.4训练生成的图像准确率为77.9%,分别超出最佳基线11.2%与15.3%。

英文摘要

Vision-Language Large Models (VLLMs) are promising for AI-generated image (AIGI) detection because they can produce both a prediction and a natural-language output. However, most existing VLLM-based detectors primarily fine-tune the language side while giving limited attention to low-level visual forensic cues. They also often depend on manually crafted prompts or human-annotated rationales, which limits scalability.We present UC-VLM, a unified multi-stage framework for AIGI detection that relies solely on binary supervision. UC-VLM first identifies effective instruction variants automatically. It then reuses the same binary label within a multi-stage training framework: (i) a visual discrimination objective that strengthens sensitivity to non-semantic forensic cues, and (ii) a label-conditioned generation objective that uses the binary label to supervise textual outputs. This design turns weak binary supervision into a shared supervision signal for both the visual pathway and the language output. Our key novelty is a unified multi-stage binary-supervised framework that consistently reuses the same authenticity labels for visual adaptation and label-conditioned text generation, while leveraging automatically optimized instructions to reduce prompt sensitivity without requiring human-written rationales or hand-crafted prompts.Experiments show that UC-VLM achieves 96.1% average accuracy on GenImage, exceeding the strongest prior result by 4.6%, and obtains 69.6% / 77.9% accuracy on Chameleon under ProGAN / SDV1.4 training, surpassing the best baseline by 11.2% / 15.3%, respectively.

发表机构

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑