arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01837cs.AI

PCSD:智能体强化学习中自蒸馏的持久一致性

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

Chunji Lv, Yangguang Wei, Junlin Liu, Yang Gao, Ming Liu, Xinming Wang, Jinyang Wu, Guoren Wang, Changsheng Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对智能体强化学习中奖励稀疏问题,提出PCSD方法,通过持久一致性推导蒸馏权重,结合GRPO优化,在ALFWorld等基准上取得优于GRPO、SDAR的性能。

中文摘要 AI 辅助

大型语言模型智能体在复杂交互任务中展现出强大潜力,但其强化学习(RL)常受稀疏奖励阻碍,因长多轮轨迹可能仅收到单一结果级信号。在线自蒸馏(OPSD)从特权教师提供密集 token 级监督,但教师在每个位置可能不可靠。现有方法通常依赖孤立的 token 级差异,易受噪声影响,或分配共享的 step 级权重,可能忽略位置变化。我们提出持久一致性自蒸馏(PCSD),其从教师偏好信号的局部持久性推导 token 级蒸馏权重。PCSD 将自适应窗口与指数衰减聚合相结合,以捕捉持久的教师相对支持,应用感知趋势的调制来衰减局部下降的支持,并通过 sigmoid 门控生成连续权重。所得目标与 GRPO 联合优化,结合密集教师指导与稀疏环境反馈。在无推理时技能的情况下,PCSD 在两个主干网络上均实现了所有基准中最佳的 ALFWorld 整体结果,较 GRPO 分别超出 15.6 和 13.3 个百分点,较 SDAR 分别超出 6.2 和 5.5 个百分点,同时在 WebShop 上保持竞争力,并在未见过的 ALFWorld 拆分上较 GRPO 提升 15.8 个百分点。

英文摘要

Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy self-distillation (OPSD) provides dense token-level supervision from a privileged teacher, but the teacher may not be reliable at every position. Existing methods commonly rely on isolated token-level discrepancies, which can be sensitive to noise, or assign a shared step-level weight that may overlook positional variation. We propose Persistent Consistency Self-Distillation (PCSD), which derives token-level distillation weights from the local persistence of teacher-favoring signals. PCSD combines adaptive windows with exponentially decayed aggregation to capture persistent relative teacher support, applies trend-aware modulation to attenuate locally declining support, and produces continuous weights through sigmoid gating. The resulting objective is jointly optimized with GRPO, combining dense teacher guidance with sparse environmental feedback. Without inference-time skills, PCSD achieves the best ALFWorld Overall results among all baselines on both backbones, exceeding GRPO by 15.6 and 13.3 points and SDAR by 6.2 and 5.5 points, while remaining competitive on WebShop and gaining 15.8 points over GRPO on unseen ALFWorld split.

↑