arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

陈旧并不意味着不安全:基础设施状态竞争下工具使用型LLM智能体的守卫精度

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao

arXiv 2609.29522首次发表:更新:

发表机构

Washington University in St. Louis; Southern Methodist University(圣路易斯华盛顿大学; 南卫理公会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对工具使用型LLM智能体在基础设施状态竞争下的安全性问题,提出用语义提交谓词守卫替代新鲜度启发式,实现高精度区分良性竞争与失效竞争,避免误阻并消除不安全提交。

AI 中文摘要

工具使用型语言模型智能体越来越多地修改调度器、数据管道、对象存储和访问控制系统。在智能体的读取与提交之间,外部状态可能发生变化,但并非每次变化都会使提交变得不安全。我们将使已声明的安全谓词失效的失效竞争与保持谓词不变的竞争及无关竞争区分开来,并探究运行时守卫能以多高的精度区分它们。我们的确定性模拟器将可见状态与权威状态分离,并在四个领域的16个基础设施任务中注入了五种非原子失效机制;冻结的智能体提案在每个控制器下以反事实方式重放,无需LLM评判。我们评估了三种提交时守卫粒度(全局纪元、读集版本、语义提交谓词)、多级验证以及模型侧门控,涉及三个本地托管的量化模型系列(Qwen3-4B、Phi-4-mini、Gemma4-8B;在单块GPU上生成3,456条轨迹)。所有三种守卫都消除了不安全提交,但它们的可用性差异显著:基于新鲜度的守卫不必要地阻止了92-95%的良性竞争,损失了高达43%的安全任务完成率,而完整的谓词守卫则未阻止任何良性竞争。这种精度依赖于契约:删除单个已声明子句会恰好将其故障类别转化为不安全提交(高达7.9%)。模型侧信号无法替代:口头置信度校准不良(ECE约为0.37),动作一致性等同于随机门控,警示提示使直接不安全率基本不变,且在新鲜度守卫阻止后,智能体会从刷新但仍不完整的读取中重新进行不安全提交。在遥测降级的情况下,隐藏的并发突变在观测上保持清洁,从而限制了所有选择性策略。因此,精确的运行时执行需要语义契约,而非新鲜度启发式或模型自我评估。

英文摘要

Tool-using language-model agents increasingly mutate schedulers, data pipelines, object stores, and access-control systems. Between an agent's read and its commit, external state can change, but not every change makes the commit unsafe. We separate invalidating races, which break a declared safety predicate, from predicate-preserving and irrelevant races, and ask how precisely runtime guards distinguish them. Our deterministic simulator separates visible from authoritative state and injects five non-atomic failure mechanisms across 16 infrastructure tasks in four domains; frozen agent proposals are replayed counterfactually under every controller without an LLM judge. We evaluate three commit-time guard granularities (global epoch, read-set version, semantic commit predicate), multi-level verification, and model-side gates on three locally hosted quantized model families (Qwen3-4B, Phi-4-mini, Gemma4-8B; 3,456 trajectories on one GPU). All three guards eliminate unsafe commits, but their availability differs sharply: freshness-based guards needlessly block 92-95% of benign races, forfeiting up to 43% of safe task completions, while the complete predicate guard blocks none. That precision is contract-dependent: deleting a single declared clause converts exactly its fault family into unsafe commits (up to 7.9%). Model-side signals do not substitute: verbal confidence is miscalibrated (ECE approximately 0.37), action agreement matches a random gate, a cautionary prompt leaves the direct unsafe rate essentially unchanged, and after a freshness-guard block agents re-commit unsafely from refreshed but still-incomplete reads. Under degraded telemetry a hidden concurrent mutation remains observationally clean, bounding every selective policy. Precise runtime enforcement therefore requires semantic contracts, not freshness heuristics or model self-assessment.

Comments8 pages, 2 figures, 11 tables. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑