发表机构
University of Chinese Academy of Sciences; Nanyang Technological University; Supply Chain Tech Team Y, JD.com; Peking University; Fudan University; Fullive-AI(中国科学院大学; 南洋理工大学; 京东Y供应链技术团队; 北京大学; 复旦大学; Fullive人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现自主LLM智能体的轨迹级安全保障存在组合性缺陷,提出非衰减循环安全状态的LoopHarness方案,可限制未授权不可逆操作数量,通过Agent-SafetyBench等评估验证其有效性。
AI 中文摘要
大型语言模型智能体正越来越多地被部署为自主循环系统。从一个人类目标出发,这类系统会反复发现任务、制定计划、执行工具调用、验证结果,并在多次无人监督的迭代中持久保存状态。然而,广泛使用的智能体安全保障措施是针对单一轨迹定义的,当下一个轨迹开始时,其安全状态会被重新初始化。我们表明,这是组合性的失败,而非实现细节问题。我们的核心结果是一种分离:针对攻击证据分散在多个迭代中的攻击,无论轨迹范围的监视器表达能力如何,其真阳性率等于假阳性率,因为它所需的证据永远不会出现在它所看到的窗口中,而保留跨迭代状态的监视器则能完美区分两者。我们进一步表明,采用几何衰减风险评分的明显修复措施是不够的,因为耐心的攻击者必须等待的冷却期是一个不随时间范围N增长的常数。随后,我们提出LoopHarness,它在循环级别恢复持久、非衰减的安全状态。在中介提交和仲裁器检测下限δ_M下,它将未授权不可逆操作的预期数量限制在B+m-1+m/δ_M,这是一个与N无关的常数,其中B+m-1项由无模型规则决定,因此在验证器完全合谋时依然有效。我们在原生Agent-SafetyBench任务上提供了完整的评估协议,包含配对的干净和被攻击情节、决定性证据仅存在于跨迭代的外部状态攻击套件、每个模块的 ablation 实验以及自适应白盒红队测试。
英文摘要
Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees, whereas a monitor retaining cross-iteration state separates the two perfectly. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon $N$. We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor $δ_M$, it bounds the expected number of unauthorized irreversible actions by $B+m-1+m/δ_M$, a constant in $N$, of which the $B+m-1$ term is decided by a model-free rule and therefore survives a fully colluding verifier. We give a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.