发表机构
State Key Laboratory of Novel Software Technology, Nanjing University; School of Intelligence Science and Technology, Nanjing University; ByteDance(南京大学计算机软件新技术国家重点实验室; 南京大学智能科学与技术学院; 字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出T2SPO方法,利用成功轨迹提供步骤级反馈,通过TabPFN估计剩余距离辅助信用分配,在ALFWorld和WebShop上提升智能体任务成功率。
AI 中文摘要
强化学习使大型语言模型(LLM)智能体能够通过与环境的交互学习多步骤行为。然而,许多交互式任务中的奖励仅反映最终结果,对哪些中间决策推进任务提供的指导有限。成功的训练轨迹包含中间状态,这些状态可以为后续交互提供监督。我们引入了轨迹到步骤策略优化(T2SPO),一种利用过去交互轨迹为策略学习提供步骤级反馈的方法。T2SPO从成功轨迹中推导出剩余距离目标,并将其与沿途访问的状态表示配对。以这些示例为条件,预训练的TabPFN回归器估计新轨迹中每个状态到成功的剩余距离。该距离估计在连续状态间的变化为智能体步骤提供辅助信用,同时结合任务级监督。随着训练的进行,新完成的轨迹刷新估计器的上下文,在不更新其参数的情况下纳入新经验。在ALFWorld和WebShop上使用1.5B和7B语言模型的实验表明,T2SPO在整体任务成功率上持续优于GRPO。
英文摘要
Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task-level supervision. As training proceeds, newly completed trajectories refresh the estimator's context, incorporating new experience without updating its parameters. Experiments with 1.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO.