arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

临界状态强化学习:诊断多轮工具使用的可训练状态

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Zixiang Chen, Wenting Zhao, Zhepeng Cen, Akshara Prabhakar, Jielin Qiu, Jianguo Zhang, Zhiwei Liu, Tulika Manoj Awalgaonkar, Liangwei Yang, Shelby Heinecke, Silvio Savarese, Huan Wang

arXiv 2609.24985首次发表:更新:

发表机构

Salesforce AI Research(赛富时人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多轮工具使用中难以定位可训练状态的问题,提出临界状态强化学习,通过嵌套采样分离动作相关奖励与噪声,并在诊断选定的状态上训练,显著提升性能,如缺失函数任务提升约14个百分点。

AI 中文摘要

多轮工具使用的失败可能取决于单次模型调用,然而仅凭奖励变化并不能揭示哪次调用能从训练中受益。当奖励依赖于后续交互时,其变化可能反映下游随机性而非当前动作之间的差异。我们引入临界状态强化学习(Critical-State RL)来识别多轮交互中的可训练状态。给定任务定义的候选调用和局部奖励,该方法评估每个奖励是否捕捉了动作对任务成功的影响,以及相对于参考策略是否存在改进空间。然后,它使用嵌套采样将动作相关的奖励变化与延续噪声分离,并在选定的状态上通过上下文赌博机训练优化策略。在伯克利函数调用排行榜(BFCL)v4上的实验比较了在诊断选定状态与替代状态进行训练的效果。对于缺失函数任务,诊断选择工具可用之后的响应;对于缺失参数任务,它选择在缺失参数被提供之前的响应。训练选定的响应能提升性能,包括在缺失函数任务上约14个百分点的提升,而训练替代响应则使性能持平或更差。我们进一步将该方法应用于多种模型和任务,包括记录的重复调用避免和内存管理。

英文摘要

Multi-turn tool-use failures can hinge on a single model call, yet reward variation alone does not reveal which call would benefit from training. When rewards depend on later interactions, their variation can reflect downstream randomness rather than differences between the current actions. We introduce Critical-State RL to identify trainable states in multi-turn interactions. Given task-defined candidate calls and local rewards, the method assesses whether each reward captures the action's effect on task success and whether improvement over a reference policy is possible. It then uses nested sampling to separate action-dependent reward variation from continuation noise and optimizes the policy at the selected states using contextual-bandit training. Experiments on the Berkeley Function Calling Leaderboard (BFCL) v4 compare training at diagnostic-selected states with training at alternative states. For missing-function tasks, the diagnostic selects the response after the tool becomes available; for missing-argument tasks, it selects the response before the missing argument is supplied. Training the selected responses improves performance, including about 14 percentage points on the missing-function task, while training the alternatives leaves performance flat or worse. We further apply the recipe across models and tasks, including logged repeat-call avoidance and memory management.

Comments31 pages, 8 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑