arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

达到还是解决?通过检查点交接归因智能体强化学习收益

Reach or Solve? Deep Diving into Agentic RL Gains with Checkpoint Handoffs

Xuan Liu, Jingbin Qian

arXiv 2609.19636首次发表:更新:

发表机构

Shanghai Jiao Tong University; Rice University(上海交通大学; 莱斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对强化学习智能体收益归因问题,提出检查点交接协议,将端到端收益分解为REACH和SOLVE,实验证明二者交互为正且可预测。

AI 中文摘要

强化学习现在训练语言模型智能体,使其在实时环境中执行数十步操作。收益是巨大的,并且这些收益被解读为更好的决策能力。闭环中的智能体自行编写其输入。每个观察都源于其自身的先前动作,因此它在回合后期遇到的状态部分是由其自身造成的。然后,在相同任务上,从不同状态对SFT检查点和RL检查点进行评分。端到端成功混合了两种变化:智能体到达的位置,以及它到达后所做的事情。将比较限制在两个策略都达到的状态并不能将它们分开。这种限制基于结果进行选择,在我们的数据中,它翻转了效应的符号。我们引入了检查点交接,这是一种评估协议,它克隆一个已发布检查点所达到的状态,并将其交给另一个检查点,无需重新训练。将到达者角色和解决者角色在SFT和RL之间交叉,将端到端收益分解为REACH和SOLVE。REACH是策略到达一个环境确认距成功还有固定动作数的状态的频率。SOLVE是它从相同的克隆状态完成的频率。在两个基准测试和两个独立发布的流程中,到达者与解决者的交互在所有五个条件下都是正的。RL历史对RL解决者的价值高于相同历史对SFT解决者的价值。在ALFWorld上,RL改善了这两个术语,并且SFT解决者在RL解决者失败的地方从未成功。独立的REACH和SOLVE差距预测了总体交互。交接仅要求一个检查点的历史可以在另一个检查点下重放,因此长视野评估可以在端到端成功之外报告到达和完成情况。

英文摘要

Reinforcement learning (RL) is widely used to improve language-model agents, and its gains are usually measured by final task success. However, an agent's earlier actions shape the states in which its later decisions are made, so final task success conflates the ability to reach useful states with the ability to complete the task once there. Comparing agents only on the states each one reaches does not separate the two, since each agent is then scored on states selected by its own actions. To address this conflation, we introduce checkpoint handoff, an evaluation protocol that decouples reaching from completing without retraining. One checkpoint acts as a reacher up to a handoff point, and another continues as the solver from the same replayed history. In detail, (1) Reach measures how often a reacher arrives at states that a replayable environment verifies to be a fixed number of actions from success, and (2) Solve measures how often a solver completes the task from identical cloned copies of those states. Crossing supervised fine-tuning (SFT) and RL checkpoints in both roles across two benchmarks and two independently released training pipelines, we find that the gain from switching the solver from SFT to RL is consistently larger when RL is the reacher, at all three model scales on TravelPlanner and on both ALFWorld splits. Further analyses on ALFWorld show that RL improves both Reach and Solve. The solver gain is larger under an RL reacher because RL reaches solvable states more often, and separately measured Reach and Solve gaps recover most of this difference. Because handoff only requires replaying one checkpoint's history under another, agentic RL evaluations can report arrival and completion alongside final success.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑