arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06984cs.CRcs.AI

HarnessSafe:评估智能体框架中持久载体的安全性

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • China Telecom Group Co., Ltd.(中国电信集团有限公司)
  • Beijing Academy of Artificial Intelligence(北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Xiao Zhang, Yusheng Wang, Yuhao Fei, Dongyuan Li, Zian Liang, Liuyu Xiang, Hongxun Gu, Zhaofeng He

AI总结:

HarnessSafe基准含7类持久载体的328个可执行案例,通过追踪攻击链的多阶段评估,揭示智能体框架中安全遏制效果的载体特异性及框架-模型配置的重要性。

AI中文摘要:

现代智能体框架(agent harnesses)通过内存、技能、工具和共享工件等持久载体在任务和会话间维持状态,但这一能力会引发延迟安全风险:受攻击者影响的内容可跨系统边界,后续影响良性请求的执行。现有基准通常聚焦于少数载体或框架,而端到端攻击成功率几乎无法体现风险的传播方式。为此,我们提出HarnessSafe,这一基准包含7类持久载体的328个可执行案例,且在多数主流智能体框架上完成评估。每个案例被定义为持久风险生命周期,追踪攻击者影响从初始进入、跨载体及系统边界持久化,到后续良性触发及可观测违规的全过程。我们进一步引入基于可观测执行证据的多阶段、追踪式评估,以确定每个攻击链的推进程度及被阻止的位置。实验显示,遏制效果具有载体特异性,且高度依赖框架-模型配置;框架和模型后端均显著影响遏制结果,而攻击成功率无法反映不同的生命周期推进模式。

英文摘要:

Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However, this capability creates delayed safety risks: attacker-influenced content can cross system boundaries and later affect the execution of a benign request. Existing benchmarks typically focus on a few carriers or harnesses, while end-to-end attack-success rates reveal little about how risks propagate. To this end, we present HarnessSafe, a benchmark comprising 328 executable cases across seven persistent-carrier families and evaluated on most mainstream agent harnesses. Each case is specified as a Persistent-Risk Lifecycle that traces attacker influence from its initial entry, through persistence across carriers and system boundaries, to a later benign trigger and an observable violation. We further introduce a multi-stage, trace-based evaluation that uses observable execution evidence to determine how far each attack chain progresses and where it is stopped. Experiments show that containment is carrier-specific and strongly depends on the harness-model configuration. Both the harness and model backend substantially shape containment outcomes, while attack success rates cannot reflect distinct lifecycle progression patterns.

补充信息

↑