arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

安全性无法组合:自主大型语言模型智能体的非衰减循环状态

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong

arXiv 2608.27141首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Nanyang Technological University; Supply Chain Tech Team Y, JD.com; Peking University; Fudan University; Fullive-AI(中国科学院大学; 南洋理工大学; 京东Y供应链技术团队; 北京大学; 复旦大学; Fullive人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现自主LLM智能体的轨迹级安全保障存在组合性缺陷,提出非衰减循环安全状态的LoopHarness方案,可限制未授权不可逆操作数量,通过Agent-SafetyBench等评估验证其有效性。

AI 中文摘要

大型语言模型智能体正越来越多地被部署为自主循环系统。从一个人类目标出发,这类系统会反复发现任务、制定计划、执行工具调用、验证结果,并在多次无人监督的迭代中持久保存状态。然而,广泛使用的智能体安全保障措施是针对单一轨迹定义的,当下一个轨迹开始时,其安全状态会被重新初始化。我们表明,这是组合性的失败,而非实现细节问题。我们的核心结果是一种分离:针对攻击证据分散在多个迭代中的攻击,无论轨迹范围的监视器表达能力如何,其真阳性率等于假阳性率,因为它所需的证据永远不会出现在它所看到的窗口中,而保留跨迭代状态的监视器则能完美区分两者。我们进一步表明,采用几何衰减风险评分的明显修复措施是不够的,因为耐心的攻击者必须等待的冷却期是一个不随时间范围N增长的常数。随后,我们提出LoopHarness,它在循环级别恢复持久、非衰减的安全状态。在中介提交和仲裁器检测下限δ_M下,它将未授权不可逆操作的预期数量限制在B+m-1+m/δ_M,这是一个与N无关的常数,其中B+m-1项由无模型规则决定,因此在验证器完全合谋时依然有效。我们在原生Agent-SafetyBench任务上提供了完整的评估协议,包含配对的干净和被攻击情节、决定性证据仅存在于跨迭代的外部状态攻击套件、每个模块的 ablation 实验以及自适应白盒红队测试。

英文摘要

Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees, whereas a monitor retaining cross-iteration state separates the two perfectly. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon $N$. We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor $δ_M$, it bounds the expected number of unauthorized irreversible actions by $B+m-1+m/δ_M$, a constant in $N$, of which the $B+m-1$ term is decided by a model-free rule and therefore survives a fully colluding verifier. We give a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑