TAPO:面向大语言模型智能体的过渡感知策略优化
TAPO: Transition-Aware Policy Optimization for LLM Agents
- School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院)
- Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出TAPO框架,通过策略优化与过渡监督交替训练,增强LLM智能体对环境过渡的敏感性,作为轻量即插即用模块,在WebShop和ALFWorld实验中提升任务性能。
AI中文摘要:
近来,强化学习(RL)已成为大语言模型(LLM)智能体后训练的关键范式。然而,现有方法主要依赖稀疏的任务奖励进行策略优化,未能充分利用在线交互过程中自然存在的另一类密集监督信号:执行动作后的环境反馈。近期理论研究表明,多步骤、面向目标任务的泛化依赖于对环境结果的预测性知识。受此启发,我们提出TAPO:面向大语言模型智能体的过渡感知策略优化,这是一种在策略优化与过渡监督之间交替进行的统一训练框架。除标准RL更新外,TAPO利用回放数据在共享主干模型上应用动作条件下的下一个观测值预测监督。该方法增强了模型对环境过渡动态和动作结果的敏感性,同时优化策略。它是现有智能体RL算法的计算轻量型即插即用增强模块,无需额外专家数据、额外采样成本或推理时间开销。我们在WebShop和ALFWorld上开展系统实验,将不同规模的基础模型与不同策略优化算法相结合。实验结果表明,TAPO相比纯策略优化基线一致提升了任务性能。
英文摘要:
Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents. However, existing methods predominantly rely on sparse task rewards for policy optimization, failing to fully exploit another class of inherently dense supervisory signals naturally present during online interaction: environmental feedback following action execution. Recent theoretical studies suggest that generalization in multi-step, goal-oriented tasks hinges on predictive knowledge of environmental consequences. Inspired by this, we propose TAPO: Transition-Aware Policy Optimization for LLM Agents, a unified training framework that alternates between policy optimization and transition supervision. Beyond standard RL updates, TAPO repurposes rollout data to apply action-conditioned next-observation prediction supervision on a shared backbone model. This approach enhances the model's sensitivity to environmental transition dynamics and action consequences while concurrently optimizing the policy. It serves as a computationally lightweight, plug-and-play enhancement module for existing agent RL algorithms, requiring no additional expert data, extra sampling costs, or inference-time overhead. We conduct systematic experiments on WebShop and ALFWorld, integrating foundation models of various scales with different policy optimization algorithms. Empirical results demonstrate that TAPO consistently improves task performance over pure policy optimization baselines.