arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03153cs.CRcs.AI

EvoRiskBench:面向工作区智能体运行时安全风险的演进式基准

EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents

  • Novo Ordo for AI(诺沃奥多人工智能机构)
  • Fudan University(复旦大学)
  • Zhongguancun Laboratory(中关村实验室)

机构由 AI 辅助整理,请以论文原文为准。

Shiyi Kuang, Xuemei Luo, Kun Liu, Junhai Li, Rui Tian, Feng Shi, Bo Shen, Nianyu Li, Dehui Li, Ping Chen

AI总结:

EvoRiskBench是一个围绕EP-Path-EF框架构建的演进式基准,通过自动化工作流在隔离环境中生成并验证450个对抗任务,评估九个模型-框架配置,发现最易受攻击配置的攻击成功率高达68.44%,揭示了工作区智能体的重大安全漏洞。

AI中文摘要:

工作区智能体将大型语言模型与执行框架相结合,以执行有状态的多步任务,这些任务会访问或修改外部资源。现有基准在运行时安全风险的可执行覆盖方面存在空白,而模型能力、框架、工具和威胁的不断演进也推动了基准的演进。我们提出了EvoRiskBench,一个围绕EP-Path-EF框架组织的演进式基准,该框架通过智能体介导的风险路径将初始风险入口点与一跳技术效应联系起来。该框架定义了九个入口点类别和五个效应类别;一项20名参与者的研究支持了其在代表性案例上的可解释性和分类一致性。在该框架的指导下,一个自动化的端到端工作流在隔离环境中构建并执行风险案例,并使用运行时轨迹和环境状态独立验证结果。该基准提供了一个可复现的数据集,包含六个场景下的450个对抗性任务。我们评估了九种模型-框架配置,涵盖三个模型(GPT-5.6 Sol、DeepSeek-V4-Pro-0813和Claude Opus 5)和三个框架(Claude Code、Codex和OpenClaw)。我们的结果揭示了各系统存在的重大漏洞。最易受攻击的配置是Codex与DeepSeek-V4-Pro-0813的组合,其攻击成功率(ASR)达到68.44%,这表明工作区智能体的配置不足以确保安全的自主执行。ASR在模型间的差异大于框架间,且框架差异取决于模型。基准案例和评估平台将在工件安全性和可复现性检查完成后发布。

英文摘要:

Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and threats motivate benchmark evolution. We introduce EvoRiskBench, an evolving benchmark organized around the EP-Path-EF framework, which links an initial risk entry point to a one-hop technical effect through an agent-mediated risk path. The framework defines nine entry-point categories and five effect categories; a 20-participant study supports their interpretability and classification consistency on representative cases. Guided by this framework, an automated end-to-end workflow constructs and executes risk cases in isolated environments and independently verifies outcomes using runtime traces and environment states. The benchmark provides a reproducible dataset of 450 adversarial tasks across six scenarios. We evaluate nine model-harness configurations spanning three models (GPT-5.6 Sol, DeepSeek-V4-Pro-0813, and Claude Opus 5) and three harnesses (Claude Code, Codex, and OpenClaw). Our results reveal substantial vulnerabilities across systems. The most vulnerable configuration, Codex with DeepSeek-V4-Pro-0813, reaches a 68.44% attack success rate (ASR), indicating that configuration of workspace agent is insufficient to ensure secure autonomous execution. ASR varies more across models than harnesses, and harness differences depend on the model. The benchmark cases and evaluation platform will be released after completion of artifact safety and reproducibility checks.

↑