AI 中文总结
研究针对LLM智能体安全评估的缺陷,推出REDAgentBench框架,经实验发现其宏平均ASR为65.69%,还揭示了识别-执行差距,无训练策略提醒可减少70%以上违规
AI 中文摘要
大型语言模型(LLM)智能体将基于语言的推理与外部工具相结合,以执行复杂任务。对抗性输入可利用智能体与其环境之间的交互,导致智能体在执行过程中违反安全策略。然而,现有评估往往将智能体安全简化为单一的攻击成功率(ASR),将暴露、执行、观察和裁决混为一谈,可能将实际违规行为与证据可见性相混淆。我们推出REDAgentBench,这是一个用于自主红队测试与忠实测量的可执行框架。它从明确的安全约束及相关的智能体系统漏洞中推导攻击,在隔离的服务沙箱中运行这些攻击,并从服务回执和最终状态变化中验证有害影响。该基准包含五个服务场景下的1661个案例。在六个模型和三个智能体 harness 上,宏平均ASR为65.69%;报告的ASR随harness和证据视图而变化,而评估上下文的披露会改变执行行为。在一个基于状态的诊断队列中,近五分之一经确认且具有已解析操作锚点的违规行为发生在智能体陈述相关约束或风险之后,这揭示了一种“识别-执行差距”。最后,一种无训练的策略提醒在匹配重放中将经确认的违规行为减少了超过70个百分点。这些发现表明,可执行评估可改进安全测量并识别可操作的干预点。
英文摘要
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAgentBench, an executable framework for autonomous red-teaming and faithful measurement. It derives attacks from explicit safety constraints and associated agent-system vulnerabilities, runs them in isolated service sandboxes, and verifies harmful effects from service receipts and final-state changes. The benchmark contains 1,661 cases across five service surfaces. Across six models and three agent harnesses, macro-average ASR is 65.69%; reported ASR varies with harness and evidence view, while evaluation-context disclosure changes execution behavior. In a state-grounded diagnostic cohort, almost one in five confirmed violations with resolved action anchors occurs after the agent states the relevant constraint or risk, revealing a Recognition--Execution Gap. Finally, a training-free policy reminder reduces confirmed violations by more than 70 percentage points in matched replay. These findings show that executable evaluation can improve safety measurement and identify actionable intervention points.
Comments6 figures, 4 tables. Supplementary material included