工具使用语言模型智能体的边界状态控制:状态漂移下的提交时一致性
Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift
- The Institute of Energetic Paradigm(能量范式研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对工具使用语言模型智能体在状态漂移下的提案到提交差距,提出BSC-R边界机制,通过绑定提交授权到动作和状态语义投影,实现提交时一致性,在多个基准上验证效果,但非普遍安全保证。
AI中文摘要:
工具使用的语言模型智能体可能在安全相关状态改变后才决定某个动作是允许的并执行它。我们研究了这种提案到提交的差距,并引入了BSC-R,一种确定性效果边界机制,它将一次性提交授权绑定到确切动作以及证明该动作合理的授权状态的语义投影。在2,847个受攻击的AgentDojo情节中,边界内核精确保留了未受保护智能体的行为(80.576%效用;1.616%攻击成功率)。在10,302个冻结提案中,它接受每个未改变的提交,并拒绝十个预先指定的改变或重放类别的每个实例。在独立生成的边界漂移实验中,完全联合绑定提交了0/4,403个无效上下文,同时保留了5,899/5,899个有效上下文。在3,460个场景的CONTINUITY套件上进行的前瞻性外部评估保留了所有700个良性案例,防止了1,200/1,200个代表性的非重放无效效果,正确处理了160/160个重放生命周期,并扣留了200/200个模糊的不发布案例。更广泛的外部套件也暴露了该方法的局限性:在所有2,560次攻击中,BSC-R的无效效果提交率为25%,而CONTINUITY为0%。因此,结果是一种有范围的提交时一致性机制,而不是普遍的智能体安全声明。
英文摘要:
Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mechanism that binds a single-use commit authorization to the exact action and to a semantic projection of the authorization state that justified it. On 2,847 attacked AgentDojo episodes, the boundary kernel preserves the unprotected agent's behavior exactly (80.576% utility; 1.616% attack success). On 10,302 frozen proposals, it accepts every unchanged commit and rejects every instance of ten prospectively specified changed or replayed classes. In an independently generated boundary-drift experiment, full joint binding commits 0/4,403 invalid contexts while retaining 5,899/5,899 valid contexts. A prospective external evaluation on the 3,460-scenario CONTINUITY suite retains all 700 benign cases, prevents 1,200/1,200 represented non-replay invalid effects, handles 160/160 replay lifecycles correctly, and withholds 200/200 ambiguous no-release cases. The broader external suite also exposes the method's limit: across all 2,560 attacks, BSC-R has a 25% invalid-effect commit rate versus 0% for CONTINUITY. The result is therefore a scoped commit-time consistency mechanism, not a universal agent-safety claim.