arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先选择后行动:面向长程工具使用智能体的比较价值估计

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li, Zheng Zhang, Xin Liu, Shengtian Yang, Guangfeng Cai, Lei Feng

arXiv 2610.02330首次发表:更新:

发表机构

Southeast University; Shandong Jianzhu University; University of South Australia(东南大学; 山东建筑大学; 南澳大利亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CITA方法,通过比较推断模型在工具调用前估计长程价值,提升长程工具使用智能体的工具F1和任务成功率。

AI 中文摘要

大型语言模型(LLMs)在复杂任务中依赖长程工具调用序列,其中每次调用都可能改变任务状态并影响后续决策。在长程工具使用中,最终结果奖励在长交互轨迹上提供的信用分配较弱。步骤级奖励能提供更有针对性的反馈,但获取可靠的步骤监督通常需要人工或LLM判断,或额外的轨迹展开来估计中间决策的下游影响。本文认为,有效的工具使用智能体应在执行可能的下一步工具调用之前估计其长程价值。这一目标需要对同一上下文下的备选调用进行对比监督,而记录的轨迹仅包含实际执行的调用。因此,我们提出面向工具使用智能体的比较推断方法(CITA)。CITA从配对信号中训练比较推断模型(CIM),这些信号结合了观察到的工具行为、来自贝叶斯工具图模拟器的可扩展监督以及基于LLM比较的语义判断。训练得到的CIM学会在当前上下文下估计可能的下一步工具调用支持最终任务成功的可能性。在三个工具使用基准和多种骨干LLM上,CITA持续提升了工具F1分数和任务成功率。额外分析表明,CIM能学习到准确的步骤级价值估计,用于比较工具选择。

英文摘要

Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same context, while logged trajectories only contain the invocation that was actually taken. Therefore, we propose Comparative Inference for Tool-use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison. The resulting CIM learns to estimate how likely a possible next tool invocation is to support final task success under the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Additional analysis shows that CIM learns accurate step-level value estimates for comparative tool choices.

CommentsNeurIPS 2026 Poster

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑