用于长期代理任务的进度条件分组策略优化
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
浏览论文内容
中文总结 AI 辅助
研究长期代理任务中基于分组策略优化的采样不平衡问题,提出进度条件分组策略优化(ProGPO),通过特定条件下利用首次访问观察覆盖率,为访问更多新状态的轨迹或步骤赋予优势,实验证明该方法优于基线,尤其在困难任务上效果显著。
中文摘要 AI 辅助
基于分组的策略优化已越来越多地用于通过比较组内轨迹或步骤,从稀疏结果奖励中训练大语言模型(LLM)智能体。然而,在困难的长期任务中,这种比较可能会出现采样不平衡:重复或低效动作在策略的高概率区域占主导,而有用的状态改变动作采样不足。这种不平衡产生许多全失败的展开组,结果奖励无法为纠正策略提供方向。这些影响共同形成自我强化的信用陷阱。为打破此循环,我们提出进度条件分组策略优化(ProGPO),仅在组内所有样本获得零结果奖励时使用首次访问观察覆盖率。具体而言,在此类组中,ProGPO为访问更多新状态的轨迹或步骤赋予更高相对优势。在两个具有挑战性的代理基准测试(ALFWorld和使用Qwen2.5 - 1.5/7B - Instruct的WebShop)上的实验表明,ProGPO始终优于基于分组的基线,在困难任务上有特别大的提升。
英文摘要
Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer from a sampling imbalance: repeated or low-effect actions dominate the high-probability region of the policy while useful state-changing actions remain under-sampled. This imbalance produces many all-failed rollout groups, where outcome rewards provide no direction for correcting the policy. Together, these effects can form a self-reinforcing credit trap: failure-dominated sampling yields no outcome-based correction, allowing repeated low-effect actions to persist. To break this loop, we propose Progress-conditioned Group Policy Optimization (ProGPO), which uses first-visit observation coverage only when all samples in a group receive zero outcome reward. Specifically, within such groups, ProGPO assigns higher relative advantages to trajectories or steps that visit more new states since reaching new observations is a prerequisite for task success. Experiments on two challenging agentic benchmarks, ALFWorld and WebShop with Qwen2.5-1.5/7B-Instruct, show that ProGPO consistently improves over group-based baselines, with particularly large gains on hard tasks.
发表机构
- Southeast University(东南大学)
- Kuaishou Technology(快手科技)
- Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。