发表机构
Tsinghua University; Zhejiang University; Meituan(清华大学; 浙江大学; 美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 AgentOPSD 这一无 critic 的递归自蒸馏方法,用于智能体强化学习的轮次级 credit 分配,在 ALFWorld 等数据集上优于 GRPO 等基线,在 Qwen2.5-7B 模型上实现 ALFWorld 89.1% 的成功率。
AI 中文摘要
带有可验证奖励的强化学习(RL)会构建轨迹级的优势估计,但在长 horizon、多轮的智能体任务中,它往往无法对决定结果的少数关键决策进行 credit 分配。近期研究引入了用于 credit 分配的特权自蒸馏,提供了更密集的监督信号,但这类局部信号应如何表示序列 credit 仍不明确。我们提出 AgentOPSD,一种无 critic 的递归方法,用于智能体强化学习中的轮次级 credit 分配。AgentOPSD 将 token 级的师生对数概率 gap 聚合为轮次级证据,并在对数几率空间中递归更新贝叶斯信念状态。这形成了一种原则性的重加权方案,将稀疏的结果监督转化为轮次级 credit 信号,并通过连续状态间的边际信念修正识别关键轮次。该方法与标准策略优化完全兼容,既不需要额外的 critic,也不需要额外的 rollouts。我们在 ALFWorld、WebShop 和 Search-QA 上使用两种规模的 Qwen2.5 模型(3B 和 7B)对 AgentOPSD 进行评估。AgentOPSD 的性能优于 GRPO 和强大的自蒸馏基线,在使用 Qwen2.5-7B 时,在 ALFWorld 上达到了 89.1% 的成功率。消融研究表明,性能提升归因于轮次级聚合和依赖历史的递归信念更新。
英文摘要
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn-level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self-distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5-7B. Ablation studies attribute the gains to turn-level aggregation and history-dependent recursive belief updates.
CommentsCode: https://github.com/ZethWang/AgentOPSD