发表机构
City University of Hong Kong; Zhejiang University; University College London; Xiaomi Corporation(香港城市大学; 浙江大学; 伦敦大学学院; 小米公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出自演进行为安全基准ActBench,结合奖励引导束搜索与双证据验证机制,评估协作智能体风险,在24000条轨迹上测试15个LLM和6个开源智能体,发现模型间攻击成功率差异更大。
AI 中文摘要
协作智能体在完成良性任务时,可能会泄露受保护数据、操纵未授权状态、调用未授权API。我们定义了行为安全,并引入ActBench——一种自演进基准,它从执行轨迹而非最终响应评估此类行为风险。每个案例配对一个良性任务与一个对抗性变体,该变体保留其指令、配置、初始状态、评分模型和可信记录,同时注入可通过任务实现的有效载荷。ActBench包含来自213个场景的600个案例,涵盖15种风险行为、6个执行空间和48个Web服务。为超越静态有效载荷,我们提出了一种奖励引导的束搜索方法,该方法联合优化攻击有效性和任务效用,同时通过反思诊断失败的执行检查点并指导有效载荷修订。此外,我们提出了一种双证据验证机制,通过日志证据和基于LLM的轨迹验证智能体执行的安全性和效用。我们在24000条轨迹上评估了15个LLM和6个开源协作智能体。在固定的测试框架下,攻击成功率在模型间为10.1%至94.4%;在固定的基础模型下,在智能体框架间为73.7%至94.4%。结果显示,模型间的差异大于智能体框架间的差异,且攻击在所有测试的智能体中仍保持高成功率。该基准已发布于:this https URL。
英文摘要
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such behavior risk from execution trajectories rather than final responses. Each case pairs a benign task with an adversarial variant that preserves its instruction, configuration, initial state, rating model, and trusted records while injecting a task-reachable payload. ActBench contains 600 cases from 213 scenarios, spanning 15 risk behaviors, six execution spaces, and 48 web-service APIs.To move beyond static payloads, we propose a reward-guided beam search method that jointly optimizes attack effectiveness and task utility, while reflection diagnoses failed execution checkpoint and guides payload revision. Besides, we propose a dual evidence verification mechanism that verifies agent execution safety and utility through log evidence and LLM-based trajectory evidence.We evaluate 15 LLMs and 6 open-source cowork agents over 24,000 trajectories. Under a fixed harness, attack success rates ranges from 10.1% to 94.4% across models, while under a fixed base model, they range from 73.7% to 94.4% across agents.These results show greater variation across models than agent harness, while attacks remain highly successful across all tested harnesses.Our benchmark is released at: https://github.com/zjuicsr/ActBench.
CommentsBenchmark and Code is available https://github.com/zjuicsr/ActBench