arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HarnessRisk:面向生命周期的智能体管控安全基准

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen

arXiv 2608.17597首次发表:更新:

发表机构

University of North Carolina at Chapel Hill; University of Central Florida; Michigan State University(北卡罗来纳大学教堂山分校; 中佛罗里达大学; 密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出面向生命周期的智能体管控安全基准HarnessRisk,含128个沙箱案例,评估多模型与管控配置的安全表现,发现管控配置阶段最易受攻击,明确风险识别未必带来安全行动,凸显多维度评估智能体安全的必要性。

AI 中文摘要

大型语言模型越来越多地通过管控工具、扩展功能、持久状态、权限和外部操作的智能体管控(agent harness)进行部署。现有的安全基准主要针对单个攻击机制或有限的操作设置子集,难以比较不同管控职责下安全故障的产生情况。我们提出了HarnessRisk,这是一个面向生命周期的基准,将智能体管控安全划分为六个操作阶段,包括管控配置(Harness Configuration)、能力扩展(Capability Extension)、运行时操作(Runtime Operation)、状态持久化(State Persistence)、操作控制(Action Control)和事件恢复(Incident Recovery)。HarnessRisk包含128个沙箱案例,每个案例将良性用户目标与嵌入在不可信任工作流工件中的对抗性指令配对。我们使用效用(Utility)、攻击成功率(Attack Success Rate)、持久化性(Persistence)和检测率(Detection)评估每个轨迹。在三个智能体管控、六个语言模型以及14种模型与管控配置的组合中,攻击成功率范围为12.6%至80.9%,而效用保持在75.0%至97.6%之间。管控配置是所有三个智能体管控中最易受攻击的阶段,表明攻击可以通过在原本授权的工作流中更改安全敏感参数来实现。我们还发现,明确的风险识别并不能可靠地导致安全行动,因为一些配置在超过90%的运行中检测到风险,但仍保留了大量攻击成功率。这些结果强调需要在多个智能体管控职责以及部署的模型和管控配置层面评估智能体安全。

英文摘要

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. HarnessRisk contains 128 sandboxed cases, each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact. We evaluate each trajectory using Utility, Attack Success Rate, Persistence, and Detection. Across three harnesses, six language models, and 14 model and harness configurations, attack success ranges from 12.6% to 80.9%, while Utility remains between 75.0% and 97.6%. Harness Configuration is the most vulnerable phase across all three harnesses, showing that attacks can succeed by altering security sensitive parameters within otherwise authorized workflows. We also find that explicit risk recognition does not reliably lead to safe action, as some configurations detect risks in more than 90% of runs while retaining substantial attack success. These results highlight the need to evaluate agent safety across multiple harness responsibilities and at the level of the deployed model and harness configuration.

CommentsProject Page: https://baiyajing.github.io/harness-risk/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑