发表机构
Beijing Jiaotong University(北京交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长 horizon LLM 智能体强化学习的信用分配难题,提出 MileGPO 方法,通过里程碑发现、RCS、PCC 三项设计提升性能,在 ALFWorld 和 WebShop 上达 SOTA 且分布差距小。
AI 中文摘要
在长 horizon 智能体强化学习中,信用分配极具挑战性,监督信号通常仅来自最终奖励。现有方法通过步骤分组或基于图的优势估计将轨迹级信号细化为步骤级信用,但可能忽略有意义的中间里程碑。我们提出 MileGPO(面向长 horizon LLM 智能体的基于图策略优化的里程碑推理),该方法通过三项设计从分组的同策略 rollout 中推导过程级信用:里程碑发现在成功 rollout 上识别候选里程碑,在失败 rollout 上识别重复陷阱;可靠性校准塑形(RCS)通过基于结果的置信度对这些候选进行加权,强化可靠的里程碑和陷阱,同时降低不确定候选的权重;进度对比校准(PCC)进一步测试候选是否反映局部进度,以及其传入状态是否优于同一状态下观察到的替代方案。该方法无需辅助模型或额外环境交互。在 ALFWorld 和 WebShop 上的实验表明,其达到了 SOTA 性能,且在 ALFWorld 上的分布内到分布外差距较小。消融实验和信用诊断表明,可靠性加权、局部进度和同状态分支证据可补充里程碑发现,解决模糊的中间信用问题。
英文摘要
Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through step grouping or graph-based advantage estimation, but can overlook meaningful intermediate milestones. We propose MileGPO (Milestone Inference with Local Evidence for Graph-Based Policy Optimization), which derives process-level credit from grouped on-policy rollouts through three designs. Milestone Discovery identifies candidate milestones on successful rollouts and recurring traps on failed ones. Reliability-Calibrated Shaping (RCS) weights these candidates by outcome-based confidence, strengthening reliable milestones and traps while down-weighting uncertain ones. Progress-Contrastive Calibration (PCC) further tests whether a candidate reflects local progress and whether its incoming transition outperforms observed alternatives from the same state. MileGPO requires neither auxiliary models nor additional environment interaction. Experiments on ALFWorld and WebShop show state-of-the-art performance and a small in-distribution to out-of-distribution gap on ALFWorld. Ablations and credit diagnostics indicate that reliability weighting, local progress, and same-state branch evidence complement milestone discovery and resolve ambiguous intermediate credit.