AI 中文总结
本文提出行动前验证框架,通过确定性检查在动作生效前捕获LLM智能体的静默失败,在shell命令和代码编辑上分别达到95.8%和99.9%的准确率,并发布基准与验证器。
AI 中文摘要
LLM智能体通过执行动作与世界交互:运行shell命令、应用代码编辑。错误的动作并不总是大声失败;它可能静默失败,产生看似合理但实际错误的效果而不引发任何错误。我们认为,在动作生效之前运行一个廉价的确定性检查,是一种有效且未被充分利用的智能体监督形式,并在一个框架下跨两种动作模态对其进行了研究。其思想是在任何执行器运行之前,通过构造固定动作的正确效果,从而直接度量静默失败,并且验证器可以选择弃权(不执行)而非猜测。对于shell命令,一个静态验证器在9930条命令和482个工具上捕获了95.8%的无效命令,误报率为10.0%。其语法和二进制检查是预言机精确的,在捕获一半错误的同时实现零误报;标志检查仅受帮助文本覆盖范围的限制,并导致了所有误报。对于代码编辑,一个包含224个文件上640个编辑的基准测试(隔离应用步骤)揭示了鲜明的差异。基于内容锚定的格式(如搜索/替换和差异)干净地失败,而基于位置锚定的格式则静默失败:行号在一行偏移下破坏了99.1%的文件,函数名编辑在12.7%的情况下命中错误的函数。在两种设置中,拒绝-当-不确定策略将静默失败转化为可恢复的失败,并以可调的适用性成本为代价:选择性接地在7.0%误报率下达到0.958的召回率,而锚定并验证的应用器在8320次试验中记录了0.01%的静默误应用。我们发布了这两个基准、验证器和防护装置。
英文摘要
An LLM agent acts on the world by emitting actions: shell commands to run, edits to apply. A wrong action does not always fail loudly; it can fail silently, producing a plausible but incorrect effect that raises no error. We argue that a cheap deterministic check, run before an action takes effect, is an effective and underused form of agent oversight, and we study it across two action modalities in one framework. The idea is to fix an action's correct effect by construction, before any executor runs, so that silent failure is measured directly and the verifier may abstain rather than guess. For shell commands, a static verifier over 9930 commands and 482 tools catches 95.8% of invalid commands at a 10.0% false-positive rate. Its syntax and binary checks are oracle-exact, giving zero false positives while catching half of all errors; the flag check is bounded only by help-text coverage and accounts for every false positive. For code edits, a benchmark of 640 edits over 224 files isolating the apply step exposes a sharp split. Content-anchored formats such as search/replace and diff fail cleanly, whereas location-anchored formats fail silently: line numbers corrupt 99.1% of files under a one-line shift, and function-name edits hit the wrong function 12.7% of the time. In both settings a refuse-when-unsure policy turns silent failures into recoverable ones at a tunable cost in applicability: selective grounding reaches 0.958 recall at 7.0% false positives, and an anchor-and-verify applier records one silent misapplication in 8320 trials (0.01%). We release both benchmarks, the verifiers, and the guards.
Comments8 pages, 3 figures