arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少些推理,多些验证:确定性门控在使用工具的语言模型智能体中恢复了一种无声的违反策略失败模式

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu

arXiv 2607.07405首次发表:更新:

发表机构

Indian Institute of Technology Kharagpur; Massachusetts Institute of Technology(印度理工学院卡拉格布尔分校; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究使用工具的语言模型智能体违反策略问题,提出用确定性预执行门控干预,在τ²基准航空公司领域评估,该方法能提高成功率,防止无声违反策略写入,虽不保证任务成功,但有可靠性成果。

AI 中文摘要

使用工具的语言模型智能体可能会违反它们被部署来执行的策略,同时看似成功完成任务。在策略宽松的环境中,工具可能会执行任何格式良好的调用,即使相应的状态转换被领域策略禁止。结果是产生一个无声的错误状态,工具和智能体的自我报告都不会暴露。我们在τ²基准航空公司领域研究这种失败模式。在一个预算有限的智能体上,78%观察到的失败是没有工具错误的无声错误状态失败,并且聚合失败率在不相交的种子上是可重现的,不是采样噪声。然后我们评估一种轻量级干预:确定性的、只读的预执行门控,在允许写入之前检查提议的调用和当前状态。一个四门套件将gpt - 4o - mini上的全基准成功率从29.6%提高到42.0%(提高12.4个百分点;配对任务级自举P = 0.0012),并且这种提升在不相交的15种子集上也能重现(提高12.3个百分点;P = 0.0008)。效果集中在门控触发的地方:在26/50个触发任务上,成功率提高了19.2个百分点,而在24个未触发任务上的变化不排除为零。两个负控制(一个自我执行的零售领域和BFCL)界定了这种机制:当工具是策略宽松时,门控有帮助,而在工具已经自我执行的地方几乎没有增加效果。作为暗示性证据,而不是核心主张,相同的失败模式在前沿仍然存在:默认推理下的gpt - 5.2仍然尝试违反策略的写入,并且相同的套件将成功率从61.2%提高到71.6%(提高10.4个百分点;P = 0.020;n = 5,无复制)。贡献是一个有界的评估和可靠性结果:确定性门控不能保证任务成功,但它们可以在动作边界确定性地防止一类已知的无声违反策略的写入。

英文摘要

Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive environments, a tool may execute any well-formed call even when the corresponding state transition is forbidden by domain policy. The result is a silent wrong state (a booking cancelled, a passenger count changed, a claim acted on without verification) that neither the tool nor the agent's self-report exposes. We study this failure mode in the $τ^2$-bench airline domain. On a budget agent, 78% of observed failures are silent wrong-state failures with no tool error, and the aggregate failure rate is reproducible across disjoint seeds, not sampling noise. We then evaluate a lightweight intervention: deterministic, read-only pre-execution gates that inspect the proposed call and current state before allowing a write. A four-gate suite raises full-benchmark success from 29.6% to 42.0% on gpt-4o-mini (+12.4pp; paired task-level bootstrap P=0.0012), and the lift reproduces on a disjoint 15-seed set (+12.3pp; P=0.0008). The effect is concentrated where the gates fire: on the 26/50 firing tasks, success rises by +19.2pp, while movement on the 24 non-firing tasks does not exclude zero. Two negative controls (a self-enforcing retail domain and BFCL) bound the mechanism: gates help when tools are policy-permissive and add little where tools already self-enforce. As suggestive evidence, not a central claim, the same failure mode persists at the frontier: gpt-5.2 at default reasoning still attempts policy-violating writes, and the same suite improves success from 61.2% to 71.6% (+10.4pp; P=0.020; n=5, no replication). The contribution is a bounded evaluation and reliability result: deterministic gates do not guarantee task success, but they can deterministically prevent a known class of silent policy-violating writes at the action boundary.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑