发表机构
UCLA; Google(加州大学洛杉矶分校; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态大语言模型在内容安全审核任务中易受攻击的问题,提出自动智能红队框架,利用多智能体架构合成对抗示例,无需人工干预,提高模型鲁棒性,降低公共图像安全基准中的误报率。
AI 中文摘要
多模态大语言模型(MLLMs)越来越多地用于细致的内容安全审核任务,但仍易受对抗攻击及分布外边缘情况影响。传统主动学习和人工标注难以应对新型多模态威胁的复杂性和数量。本文提出自动智能红队框架,用迭代策略系统合成困难示例,利用多智能体架构,无需人工干预自主发现违规和边缘情况。通过将合成的对抗示例用作测试时检索的上下文示范,显著提高目标模型鲁棒性,在公共图像安全基准中,将误报率从41.2%降至24.5%。
英文摘要
Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases. Traditional active learning and manual annotation fail to scale against the complexity and volume of novel multimodal threats. In this paper, we propose an automated, agentic red-teaming framework that systematically synthesizes difficult examples using an iterative strategy that proposes novel hypotheses as well as mutating on past attempts. Leveraging a multi-agent architecture that consists of a high-reasoning Architect agent, an advanced image generator, and a multi-level verification committee of LLM raters, our system autonomously uncovers boundary-pushing violations and ambiguous policy edge cases without any human intervention. By employing these carefully synthesized adversarial examples as in-context demonstrations via test-time Retrieval, we substantially improve the target model's robustness, reducing the False Negative Rate (FNR) from 41.2% to 24.5% in a public image safety benchmark without relying on any human labeling.
Comments23 pages; work in progress