arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21027cs.AI

无需解决,仅需比较:用于LLM智能体运行时干预的微型顾问

Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents

Yanze Jiang, Mingxuan Li, Yuhao Wang, Shengfang Zhai, Jiaheng Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM智能体运行时干预需求,提出仅比较框架COTA,经多任务多执行器评估,其性能优于基线,可在辅助模型能力弱于执行器时实现有效干预。

中文摘要 AI 辅助

LLM智能体正成为需要推理、工具使用和序列决策的现实世界任务的重要范式。随着这些智能体的运行周期变长,运行时干预提供了一种无需重新训练底层执行器即可提高可靠性的方法。仅故障检测是不够的,有效的干预还必须提供有用的恢复方向。现有方法通常依赖于专家求解器或生成特定任务修正的评判者,这要么产生另一个有能力的求解器的成本,要么产生任务型评判者的容量需求。我们提出了仅比较框架COTA(Comparison-Only Tiny Advisor,仅比较微型顾问),用于建设性运行时干预。在COTA中,一个微型比较器判断采样的替代方案是否比执行器的提议带来更好的后续结果,重复的比较确定何时需要干预。我们使用由相同前缀反事实分支构建的成对监督来训练比较器,将偏好的替代方案作为非约束性建议返回,让原始执行器重新规划。在使用三个执行器的WebShop、ALFWorld和tau^3-Retail上,COTA改进了所有9个评估设置,并优于所有对比基线。这些结果表明,即使辅助模型的任务解决能力远弱于执行器,建设性运行时干预仍能保持有效。

英文摘要

LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve reliability without retraining the underlying actor. Effective intervention must provide a useful direction for recovery besides a warning. Existing approaches often rely on an expert solver or a critic that generates task-specific corrections, incurring either the cost of another capable solver or the capacity demands of a task-capable critic. We introduce Comparison-Only Tiny Advisor (COTA) for constructive runtime intervention, which reduces the learned intervention role to local action comparison. A lightweight comparator judges the actor's proposal against available alternatives, and preferred alternatives are returned as non-binding advice for replanning. The comparator is trained from same-prefix counterfactual branches. Across WebShop, ALFWorld, and tau^3-Retail with three LLM actors, COTA instantiated with a 0.5B comparator consistently improves the original actor and achieves the strongest overall performance--cost trade-off among the compared methods. These results suggest that effective runtime intervention need not itself be a task-solving problem: the intervention role can be separated from task solving and handled by a lightweight model specialized for local comparison.

发表机构

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑