arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06094cs.CV

文本到图像模型隐式漏洞的自动红队测试

Automatic Red Teaming for Implicit Vulnerabilities of Text-to-Image Models

  • Columbia University(哥伦比亚大学)
  • University of Oxford(牛津大学)
  • LMU Munich(慕尼黑大学)
  • Siemens AG(西门子股份公司)
  • Konrad Zuse School of Excellence in Reliable AI (relAI)(康拉德·楚泽卓越可靠人工智能学院(relAI))

机构由 AI 辅助整理,请以论文原文为准。

Chang Ma, Junlin Han, Shuo Chen, Runjia Li, Philip Torr, Jindong Gu

AI总结:

针对文本到图像模型难以检测的隐式对抗提示,提出多模态智能体框架AdvPIE,通过策略与评判智能体协作及累积对抗性解码,无需参数访问即可有效暴露隐式漏洞并优于基线。

AI中文摘要:

对文本到图像(T2I)模型进行红队测试对于安全部署至关重要,然而,针对隐式对抗性提示,这一任务尤为艰巨。与那些容易被识别和拦截的显式对抗性提示不同,隐式对抗性提示更难被检测:这些提示在文本表面看似良性,却仍能导致不适当的视觉内容生成。为解决这一问题,我们提出了隐式漏洞对抗性探测(AdvPIE),这是一种多模态智能体框架,旨在无需访问目标模型参数的情况下暴露隐式漏洞。AdvPIE采用一个策略智能体,根据评判智能体的反馈来生成并优化隐式对抗性提示。为了构建信息丰富的反馈,评判智能体在多次迭代中提供全局和相对两个层面的模态特定安全评估。为了有效利用反馈,我们为策略智能体提出了一种新颖的累积对抗性解码策略,该策略动态地重新加权令牌分布,以偏向那些能导致更有害图像生成的令牌,同时保持采样多样性。在标准和安全对齐的T2I模型上进行的大量实验表明,AdvPIE能有效揭示隐式漏洞,性能优于多种基线方法。

英文摘要:

Red-teaming Text-to-Image (T2I) models is essential for safe deployment, yet it remains particularly challenging against implicit adversarial prompts. Unlike explicit adversarial prompts that can be readily identified and blocked, implicit ones are much harder to detect: the prompts appear benign on the text surface yet still lead to inappropriate visual content. To address this, we propose Adversarial Probing for Implicit VulnErabilities (AdvPIE), a multimodal agentic framework to expose implicit vulnerabilities without requiring access to the parameters of target models. AdvPIE adopts a policy agent to generate and refine implicit adversarial prompts based on the feedback from a judge agent. To construct informative feedback, the judge agent provides modality-specific safety evaluation at both global and relative levels across iterations. To effectively leverage the feedback, we propose a novel Cumulative Adversarial Decoding strategy for the policy agent, which dynamically reweights token distributions to favor tokens that lead to more harmful images while preserving sampling diversity. Extensive experiments on standard and safety-aligned T2I models show that AdvPIE1 effectively uncovers implicit vulnerabilities, outperforming various baseline methods.

补充信息

↑