发表机构
Xiangtan University; Jinan University; Huazhong University of Science and Technology; City University of Hong Kong; Changsha University of Science and Technology; Zhejiang University; Chongqing University(湘潭大学; 暨南大学; 华中科技大学; 香港城市大学; 长沙理工大学; 浙江大学; 重庆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文推出首个评估商业图像生成器视觉虚假信息风险的基准EpiReal-Bench,结合EpiReal-Attack框架,发现多数商业模型易生成虚假主张的可信图像,对齐机制存在盲点。
AI 中文摘要
图像生成模型如今能够产出包含丰富文本、外观逼真的视觉产物,这类产物难以与新闻报道、教科书页面等真实世界证据区分。但同一能力也带来了新风险:这些模型能轻易生成视觉虚假信息,就连GPT-Image-2这类商业模型也能轻松生成。值得注意的是,研究发现这些模型被要求判断主张是否虚假时能识别出虚假主张,却仍会将该主张渲染为可信的视觉证据。这种差异表明当前对齐机制存在盲点:安全措施仅判断图像展示的内容,而非图像所断言的内容;然而现有的红队测试基准针对的是暴力或露图像等常规有害内容,对视觉虚假信息的对齐边界,尤其是商业模型的相关情况鲜有涉及。为填补这一空白,本文推出EpiReal-Bench,这是首个用于评估商业图像生成器视觉虚假信息风险的系统性基准,包含1万条虚假主张提示和1万张对应生成图像,覆盖10类真实世界主张和10种可信视觉格式。本文还推出EpiReal-Attack,这是一种技能引导的黑盒优化框架,采用基于帕累托的选择和多模态反馈来识别能绕过对齐安全措施,同时保持视觉真实性、文本可读性和语义保真度的指令。对4种商业模型的实验显示,超过70%的虚假主张提示会触发准确描绘对应虚假信息的图像,而EpiReal-Attack将这一比例推至95%。最令人担忧的是,这些模型仅需一次点击就能使用,其输出传播成本低却难以被质疑,使得这一维度的对齐机制基本未受保护。
英文摘要
Image-generation models can now produce text-rich, natural-looking visual artifacts that are hard to distinguish from real-world evidence, such as news reports and textbook pages. Yet, the same capability introduces a new risk: these models can just as easily fabricate visual misinformation. Even commercial models (e.g., GPT-Image-2) readily produce it. Curiously, we find that these models can recognize a claim as false when asked, yet still render that very claim as credible visual evidence. This discrepancy points to a blind spot in current alignment: safeguards judge what an image shows, not what it asserts; however, existing red-teaming benchmarks target conventional harmful content, such as violent or explicit imagery, and say little about where the alignment boundaries lie for visual misinformation, especially in commercial models. To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats. We further introduce EpiReal-Attack, a skill-guided black-box optimization framework that uses Pareto-based selection and multimodal feedback to identify commands that bypass alignment safeguards while preserving visual realism, textual legibility, and semantic fidelity. Experiments on four commercial models reveal that more than 70% of false-claim prompts elicit images that faithfully depict the corresponding misinformation, and EpiReal-Attack pushes this rate to 95%. Most worryingly, these models are only a click away, and their outputs are cheap to spread yet hard to disbelieve, leaving this dimension of alignment largely unguarded.
Comments29 pages, 19 figures, 6 tables. Project website: https://github.com/Ye-ze-yu/EpiReal-Bench