发表机构
Columbia University; Zoom Communications, Inc.(哥伦比亚大学; Zoom Communications, Inc.)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RELACE 通过回溯似然比较和状态条件优势估计,为长时程语言智能体提供细粒度动作信用,无需额外模型,在 ALFWorld 和 WebShop 上显著优于 GRPO、GiGPO 和 HCAPO。
AI 中文摘要
组相对策略优化(GRPO)通过从 rollout 组中估计优势来避免使用独立的评论家。然而,对于多轮智能体,轨迹级监督提供了粗糙且嘈杂的信用:终端奖励不能定位错误,并且可能将有用动作与错误一起惩罚。组内组策略优化(GiGPO)及后续方法通过状态条件比较来细化监督,但其信用估计仍对下游决策和结果敏感。我们引入了 RELACE(Retrospective Likelihood-based Action credit Estimation),一个无评论家框架,将回溯动作评估与状态条件优势估计相结合。RELACE 通过教师强制似然评分在原始上下文和结果增强上下文下评估已执行的动作。比较这些似然产生一个轨迹归一化的回溯因子,该因子捕获动作合理性的结果依赖变化,而不仅仅是事后合理性。我们使用该因子对折扣任务回报进行重新加权,并通过比较任务中等价状态的动作的加权回报来构建局部优势。这将回溯相关性与观测奖励耦合,产生细粒度信用,补充轨迹级 GRPO 监督。时间平滑和成功保护掩码进一步稳定局部信号。RELACE 不需要辅助价值或奖励模型,也不需要额外的自回归 rollout 进行信用估计。在 ALFWorld 和 WebShop 上使用 Qwen2.5-1.5B-Instruct 和 Qwen2.5-7B-Instruct 的实验表明,与 GRPO、GiGPO 和 HCAPO 相比有显著改进。使用 1.5B 模型,RELACE 在 ALFWorld 上达到 96.35% 的成功率,在 WebShop 上达到 79.43%,分别超过 GiGPO 5.47 和 5.60 个百分点。
英文摘要
Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain sensitive to downstream decisions and outcomes. We introduce RELACE, Retrospective Likelihood-based Action, a critic-free framework that integrates retrospective action assessment with state-conditioned advantage estimation. RELACE evaluates executed actions through teacher-forced likelihood scoring under both their original contexts and outcome-augmented contexts. Comparing these likelihoods yields a trajectory-normalized retrospective factor that captures outcome-dependent changes in action plausibility, rather than hindsight plausibility alone. We use this factor to reweight discounted task returns and construct local advantages by comparing weighted returns among actions from equivalent states within a task. This couples retrospective relevance with observed reward, producing fine-grained credit that complements trajectory-level GRPO supervision. Temporal smoothing and success-protecting masking further stabilize the local signal. RELACE requires neither auxiliary value nor reward models nor additional autoregressive rollouts for credit estimation. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct demonstrate substantial improvements over GRPO, GiGPO, and HCAPO. With the 1.5B model, RELACE achieves $96.35\%$ success on ALFWorld and $79.43\%$ on WebShop, surpassing GiGPO by $5.47$ and $5.60$ percentage points, respectively.