PlanPO:面向多轮智能体大语言模型的组规划感知策略优化
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs
AI总结:
针对现有组相对策略优化无法区分成功轨迹效率差异的问题,提出 PlanPO 方法,引入由粗到细的优势信号,在三类多轮基准上较 GRPO 平均提升 27.2%,且训练成本可忽略。
AI中文摘要:
组相对策略优化已成为在多轮交互任务上训练智能体大语言模型(LLMs)的关键范式。然而,现有大多数变体无法区分成功轨迹间的优势,即便这些轨迹在交互效率上存在显著差异。例如,迂回的成功轨迹常被赋予相同的结果奖励,导致优势崩溃和严重的性能瓶颈。为此,我们提出组规划感知策略优化(PlanPO),这是一种简单却有效的强化学习方法,用于学习超越任务特定高质量行为模式的可泛化规划能力。具体而言,PlanPO引入了由粗到细的优势信号,该信号捕获针对同一任务采样的成功轨迹在轨迹级长度和轮次级响应长度上的相对差异。在组相对优化结构内,这使得智能体能够从高质量 rollout 中主动学习涵盖交互规划和文本生成的可泛化且审慎的行为,而不会退化为普通的长度最小化。实验表明,在具有挑战性的多轮基准 ALFWorld、WebShop 和 SciWorld 上,PlanPO 平均比 GRPO 提升了 27.2%,优于近期强大的基线,同时仅产生可忽略的额外训练成本。
英文摘要:
Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories differ substantially in their interaction efficiency. For instance, circuitous successes are often assigned the identical outcome reward, causing advantage collapse and severe performance bottlenecks. To this end, we propose Group Planning-aware Policy Optimization (PlanPO), a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns. Specifically, PlanPO introduces coarse-to-fine advantage signals, which capture the relative differences in trajectory-level lengths and turn-level response lengths conditioned on successful trajectories sampled for the same task. Within the group-relative optimization structure, this enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generation from high-quality rollouts, without degenerating into vanilla length minimization. Experimentally, PlanPO improves over GRPO by 27.2\% on average across the challenging multi-turn benchmarks ALFWorld, WebShop, and SciWorld, outperforming recent powerful baselines while incurring negligible additional training cost.