arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向智能体强化学习的组内自蒸馏方法

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

Binbin Zheng, Zijun Xie, Guanqun Zhao, Enlei Gong, Xing Ma, Xiaoliang Fu, Zeyu Chen

arXiv 2607.28076首次发表:更新:

发表机构

Baidu(百度)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对智能体强化学习中终端奖励监督粗粒度的问题,提出组内自蒸馏方法,利用策略自身已验证轨迹生成指导,细化回合级信用分配,在多环境多模型上性能优于基线且泛化性更好。

AI 中文摘要

带可验证奖励的强化学习(RLVR)可有效训练大语言模型智能体,但终端奖励仅提供粗粒度的轨迹级监督,导致成功行为、重复错误和偶然选择在同一结果信号中混杂。现有智能体自蒸馏方法用自然语言技能丰富稀疏监督,但外部检索或由更强模型从单条轨迹提取的技能可能与当前经验不匹配、超出策略能力或仅适用于特定路径。本文提出组内自蒸馏(GRSD),从策略自身的已验证 rollout 中推导与能力对齐、结果可区分的指导:对每个提示,策略在同策略组内反思每条已验证轨迹,停止梯度快照对比成功与失败 rollout 的反思结果,构建组级特权指导;基于该指导,自教师通过调节基于结果的优势来细化回合级信用分配,同时保留验证器确定的学习方向。在多个智能体环境和模型规模上的实验表明,GRSD 始终优于竞争性基线,且对未见任务的泛化效果更好。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only coarse trajectory-level supervision, leaving successful behaviors, recurring mistakes, and incidental choices entangled in the same outcome signal. Existing agentic self-distillation methods enrich sparse supervision with natural-language skills, but skills retrieved externally or extracted from a single trajectory by stronger models may mismatch current experience, exceed the policy's capability, or remain path-specific. We propose Group-Reflective Self-Distillation (GRSD), which derives capability-aligned and outcome-discriminative guidance from the policy's own verified rollouts. For each prompt, the policy reflects on each verified trajectory in an on-policy group, and a stop-gradient snapshot contrasts the resulting reflections from successful and failed rollouts to construct group-level privileged guidance. Conditioned on this guidance, a self-teacher refines turn-level credit assignment by modulating outcome-based advantages while preserving the verifier-determined learning direction. Experiments across multiple agentic environments and model scales demonstrate that GRSD consistently outperforms competitive baselines and generalizes more effectively to unseen tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑