arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27336cs.AI

CART:大型语言模型的闭环自适应红队测试

CART: Closed-Loop Adaptive Red Teaming for Large Language Models

Dongdong Zhang, Tengchao Lv, Yilin Jia, Yuzhong Zhao, Yupan Huang, Wenshan Wu, Xiangyang Zhou, Shaohan Huang, Nan Yang, Li Dong, Lei Cui, Furu Wei

首次发表
浏览论文内容

中文总结 AI 辅助

CART提出闭环自适应红队测试框架,通过挑战者、目标与评判者分离,动态引导测试,在多个评估系列中发现更多失败和更高平均风险,将红队测试转为持续可审计的弱点搜索。

中文摘要 AI 辅助

自动化红队测试通常重放一组固定的提示,这只能衡量已知风险,而无法从测试过程中发现的失败中学习。我们提出了CART(闭环自适应红队测试),这是一个利用每次测试结果来指导下一步测试内容的框架。CART从广泛的风险覆盖开始,跟踪出现的弱点,保持新探测的多样性,并记录每项发现的证据和来源。它将创建测试的挑战者、被测试的目标(可能是纯文本模型或受限的工具使用智能体)以及评估结果的评判者分离开来,使得这些角色可以被独立研究。在三个评估系列(Frontier、JAH和Agentic)中,对于每个具有可用基线的目标,CART发现的失败次数和平均风险均高于静态种子重放。这些增益扩展到工具介导的智能体测试,表明上下文适应可以揭示直接提示重放无法触及的弱点。这些结果描述了测试策略所发现的,而非实际部署中失败发生的频率。我们还发现,挑战者-评判者的选择会影响所揭示的证据,凸显了角色分离和独立审查的必要性。总体而言,CART将红队测试从一次性的清单转变为对模型和智能体弱点的持续、自适应且可审计的搜索。

英文摘要

Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next. CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding. It separates the Challenger that creates tests, the Target being tested, which may be a text-only model or a bounded tool-using agent, and the Judge that evaluates the results, allowing these roles to be studied independently. Across three evaluation families (Frontier, JAH, and Agentic), CART discovers more failures and higher average risk than static seed replay for every Target with an available baseline. The gains extend to tool-mediated agent tests, suggesting that contextual adaptation can reveal weaknesses that direct prompt replay does not exercise. These results describe what the test policies discover, not how often failures occur in real deployments. We also find that Challenger-Judge choices affect the evidence uncovered, highlighting the need for role separation and independent review. Overall, CART turns red teaming from a one-time checklist into a continuous, adaptive, and auditable search for model and agent weaknesses.

发表机构

  • Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑