arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00677cs.CL

OpenART:通过开放式环境演化扩展智能体红队演练规模

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

发表机构复旦大学 · 上海人工智能实验室 · XSafeAI
查看机构详情
  • Fudan University(复旦大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • XSafeAI

机构由 AI 辅助整理,请以论文原文为准。

Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu, Jie Li, Yan Teng, Xingjun Ma, Xia Hu, Yu-Gang Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对现有智能体安全基准无法捕捉累积风险的问题,推出开放式智能体红队演练平台OpenART,并提出EMHA攻击方法,在75种智能体模型配置上实现85.0%的汇总攻击成功率,为复杂环境下的智能体安全研究提供可扩展基础。

中文摘要 AI 辅助

AI智能体在持续运行的环境中工作,早期的状态变化会对后续的决策产生深远影响。与传统的语言模型交互不同,智能体的行为由共享状态介导,该状态会在长周期工作流中被反复修改和复用。当前的安全基准往往无法捕捉这些累积风险,因为它们聚焦于简短、静态的任务。为解决这些局限,我们推出OpenART,这是一个通过环境演化实现可扩展智能体红队演练的开放式平台。OpenART涵盖50个领域的10000多个经过验证的有状态场景,依托超过500000种工具和技能池构建。这些任务平均需要97次工具调用,可对75种不同的智能体模型配置进行统一评估。为系统探索这些不断演化的攻击面,我们提出演化马尔可夫超图攻击(Evolutionary Markov Hypergraph Attack,EMHA)。EMHA是一种黑盒策略,通过协调授权的状态转换执行反馈驱动的环境演化,无需参数更新。在整个评估过程中,任务目标保持固定,仅环境状态发生变化。在所有配置中,EMHA的汇总攻击成功率(Attack Success Rate,ASR)达到85.0%。它相较于仅基于指令的演化的优势,在简单环境中约为2%,在最复杂的环境中则超过17%,表明随着任务复杂度提升,环境演化能更有效地暴露安全缺陷。此外,我们的分析显示,除了底层模型的能力外,智能体的具体运行时实现也能解释安全性能差异的很大一部分。这些结果确立了OpenART作为在复杂、演化环境中研究智能体安全的可扩展基础的地位。

英文摘要

AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, static tasks. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART provides over 10,000 validated stateful scenarios across 50 domains, drawing from a pool of more than 500,000 tools and skills. These tasks require a median of 97 tool calls and enable unified evaluation across 75 different agent-model configurations. To systematically explore these evolving attack surfaces, we propose the Evolutionary Markov Hypergraph Attack (EMHA). EMHA is a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates. Throughout the evaluation, task objectives remain fixed while only the environment state changes. Across all configurations, EMHA achieves a pooled Attack Success Rate (ASR) of 85.0%. Its advantage over instruction-only evolution increases from approximately 2% on simple environments to over 17% on the most complex ones, demonstrating that environment evolution increasingly exposes safety failures as task complexity grows. Furthermore, our analysis shows that the specific runtime implementation of an agent explains a significant portion of safety variation beyond the underlying model's capabilities. These results establish OpenART as a scalable foundation for studying agent safety in complex, evolving environments. Code is avaible: https://github.com/AI45Lab/OpenART#

↑