智能体何时可以停止?带证据的工具使用大语言模型终止机制
When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs
浏览论文内容
中文总结 AI 辅助
该研究针对工具使用大语言模型的终止问题,提出带证据的终止(ECT)机制,经实验验证其能显著减少不安全完成与过早无支持终止,满足非劣效性要求,可实现成功恢复。
中文摘要 AI 辅助
使用工具的智能体必须决定何时停止。现有系统已能控制终端成功、验证执行轨迹或强制运行时策略,但未在受控终止故障下于COMPLETE边界测试这种特定的收据、范围和封闭重放设计。我们实例化并评估带证据的终止(Evidence-Carrying Termination, ECT):仅当智能体返回类型化证书,将每个所需答案声明绑定到有效、在范围内的轨迹证据,且确定性重放重构所声明的值时,智能体才可返回COMPLETE。一项锁定的静态研究涵盖6个工具使用类别中的48个完全合成任务,包含干净执行和8种故障。ECT产生0/288个不安全完成,而被检查的终止评判核心则为252/288(差值-87.50个百分点,95%任务集群区间[-87.50, -87.50]个百分点)。随后一项预先指定且冻结的含576条轨迹的新研究,将ECT与评判核心、其忠实控制器及全轨迹大语言模型评判器进行比较。在22个主要保留任务集群上,ECT产生0/66个过早的无支持终止,而控制器为40/66(差值-60.61个百分点,95%区间[-78.79, -40.91]个百分点);同时支持的完成率为97/132对92/132(差值3.79个百分点,区间[0.00, 9.09]个百分点),满足-10个百分点的非劣效性边界。ECT在66条轨迹中实现18次成功恢复,其中17次随后获得支持完成;所有三个闭环门均通过。ECT在声明的假设下于记录的轨迹中验证支持,而非外部真相、安全性或对齐。
英文摘要
Tool-using agents must decide when to stop. Existing systems already gate terminal success, certify execution traces, or enforce runtime polici es, but do not test this particular receipt-, scope-, and closed-replay design at the COMPLETE boundary across controlled termination faults. W e instantiate and evaluate Evidence-Carrying Termination (ECT): an agent may return COMPLETE only when a typed certificate binds every required answer claim to valid, in-scope trace evidence and a deterministic replay reconstructs the claimed value. A locked static study crosses 48 ful ly synthetic tasks in six tool-use families with clean execution and eight faults. ECT produced 0/288 unsafe completions versus 252/288 for the inspected termination-critic core (difference -87.50 pp, 95% task-cluster interval [-87.50, -87.50] pp). A fresh, prespecified and frozen 576- trajectory study then compares ECT with the critic core, its faithful controller, and a full-trace LLM critic. On 22 primary held-out task clus ters, ECT produced 0/66 premature unsupported terminations versus 40/66 for the controller (difference -60.61 pp, 95% interval [-78.79, -40.91] pp), while supported completion was 97/132 versus 92/132 (difference 3.79 pp, interval [0.00, 9.09] pp), satisfying a -10-point noninferiority margin. ECT executed successful recovery in 18/66 trajectories, of which 17 subsequently completed with support; all three closed-loop gates p assed. ECT certifies support in a recorded trace under declared assumptions, not external truth, safety, or alignment.
发表机构
- University of California San Diego(加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。