arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26048cs.AIcs.CLcs.SE

FIRE:面向可靠语言模型智能体的失败知情运行时工程

FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents

  • Failproof AI

机构由 AI 辅助整理,请以论文原文为准。

Nikita Agarwal, Nivedit Jain

AI总结:

本研究提出FIRE运行时策略,通过自然语言指令和动作拒绝在不改变模型权重的情况下,将语言模型智能体的重复成功率提升最多9.2个百分点,显著增强其交付可靠性。

AI中文摘要:

语言模型智能体常常能达到一个可行的解决方案,但随后却无法稳定地交付该方案。我们研究运行时策略:由智能体框架在观察到失败之前的状态下应用的有针对性的自然语言指令和动作拒绝,而不改变模型权重或用户提示。通过这种方式,在保持能力不变的情况下,我们观察到交付可靠性的显著提升。在完整的包含87个任务的Terminal-Bench 2.1套件中,每个任务尝试两次,策略在所有三个GPT-5.6层级中均提高了重复成功(pass^2):Luna从50.6%提升至54.0%,Terra从55.2%提升至60.9%,Sol从64.4%提升至73.6%。Sol的最佳两次尝试成功率变化了1.2个百分点,而重复成功率上升了9.2个百分点,这表明策略主要将可达成的解决方案转化为可靠的交付。我们进一步在Terra的冻结组合下覆盖了14个任务。策略引导的Terra达到了71.4%,而未辅助的Sol为64.3%,且成本约为其一半,展示了围绕模型的工程如何为用例解锁可靠性。为了隔离机制,我们进行了一项随机五臂实验:真实策略在符合条件的任务上达到61%,而无策略时为39%,时间匹配的假策略时为36%,通用验证或重新考虑时为39%至43%。预期的纠正行为出现在24个编码策略尝试中的22个,而任何其他臂中最多为14个。因此,运行时策略是一个实用的可靠性层:它们使智能体已经拥有的能力变得显著更可重复。

英文摘要:

Language-model agents often reach a working solution and then fail to consistently deliver it. We study runtime policies: targeted natural-language instructions and action denials applied by the agent harness at states that preceded observed failures, without changing model weights or the user prompt. With this, keeping capability constant, we observe a meaningful unlock in delivered reliability. Across the complete 87-task Terminal-Bench 2.1 suite, with two attempts per task, policies increase repeated success (pass^2) in all three GPT-5.6 tiers: 50.6% to 54.0% for Luna, 55.2% to 60.9% for Terra, and 64.4% to 73.6% for Sol. Sol's best-of-two success changes by 1.2 points while repeated success rises by 9.2, showing that policies chiefly convert reachable solutions into dependable delivery. We further cover 14 tasks under Terra's frozen portfolio. Policy-guided Terra reaches 71.4%, compared with 64.3% for unassisted Sol, at about half the cost, demonstrating how engineering around models could unlock dependability for a use case. To isolate the mechanism we run a randomized five-arm experiment: real policies reach 61% on eligible tasks, versus 39% without a policy, 36% with a timing-matched sham, and 39 to 43% with generic verification or reconsideration. The intended corrective behavior appears in 22 of 24 coded policy attempts, against at most 14 in any other arm. Runtime policies are therefore a practical reliability layer: they make capabilities an agent already possesses substantially more repeatable.

↑