arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于多级智能数据管理的自动硬示例合成

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation

Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan, Nichole J. Hansen, Blaž Bratanič, Nathan L Clement, Shalini Ghosh, Ariel Fuxman

arXiv 2607.14256首次发表:更新:

发表机构

UCLA; Google(加州大学洛杉矶分校; 谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态大语言模型在内容安全审核任务中易受攻击的问题,提出自动智能红队框架,利用多智能体架构合成对抗示例,无需人工干预,提高模型鲁棒性,降低公共图像安全基准中的误报率。

AI 中文摘要

多模态大语言模型(MLLMs)越来越多地用于细致的内容安全审核任务,但仍易受对抗攻击及分布外边缘情况影响。传统主动学习和人工标注难以应对新型多模态威胁的复杂性和数量。本文提出自动智能红队框架,用迭代策略系统合成困难示例,利用多智能体架构,无需人工干预自主发现违规和边缘情况。通过将合成的对抗示例用作测试时检索的上下文示范,显著提高目标模型鲁棒性,在公共图像安全基准中,将误报率从41.2%降至24.5%。

英文摘要

Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases. Traditional active learning and manual annotation fail to scale against the complexity and volume of novel multimodal threats. In this paper, we propose an automated, agentic red-teaming framework that systematically synthesizes difficult examples using an iterative strategy that proposes novel hypotheses as well as mutating on past attempts. Leveraging a multi-agent architecture that consists of a high-reasoning Architect agent, an advanced image generator, and a multi-level verification committee of LLM raters, our system autonomously uncovers boundary-pushing violations and ambiguous policy edge cases without any human intervention. By employing these carefully synthesized adversarial examples as in-context demonstrations via test-time Retrieval, we substantially improve the target model's robustness, reducing the False Negative Rate (FNR) from 41.2% to 24.5% in a public image safety benchmark without relying on any human labeling.

Comments23 pages; work in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑