arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14777cs.CL

SEED:用于智能体强化学习的自进化在线策略蒸馏

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

  • Tsinghua University(清华大学)
  • Zhejiang University(浙江大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • Nanyang Technological University(南洋理工大学)
  • Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao

AI总结:

研究针对基于结果的强化学习中间决策指导有限的问题,提出SEED框架。它将在线策略轨迹转化为训练技能并提炼回策略模型,通过微调策略生成技能,重新评分动作转化蒸馏信号,联合优化提升性能与样本效率及泛化能力。

AI中文摘要:

大型语言模型越来越多地被训练为用于涉及多轮交互、工具使用和环境反馈的长期任务的交互式智能体。基于结果的强化学习(RL)提供了一种实用的优化范式,但其稀疏的轨迹级奖励对中间决策的指导有限,在情节级结果和令牌级策略学习之间存在监督差距。我们提出了SEED(自进化在线策略蒸馏),这是一个自进化框架,将完成的在线策略轨迹转换为训练时的事后诸葛亮技能,并将其行为效果提炼回策略模型。SEED首先微调策略以分析完成的轨迹并生成捕获可重复使用工作流程、决定性观察或避免失败规则的自然语言技能。在RL期间,当前策略既收集轨迹,又作为从中提取事后诸葛亮技能的分析器。因此,策略更新共同改进后续决策和技能分析,使事后诸葛亮监督随策略一起发展。然后,SEED在普通和技能增强的上下文中对采样动作重新评分,将技能引起的概率转移转换为密集的令牌级在线策略蒸馏信号。该信号与基于结果的RL联合优化,使辅助监督与当前轨迹分布保持一致。在基于文本和基于视觉的智能体任务上的大量实验表明,SEED持续提高性能和样本效率,对未见场景具有强大的泛化能力。我们的代码可在这个https网址获取。

英文摘要:

Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on intermediate decisions, leaving a supervision gap between episode-level outcomes and token-level policy learning. We propose SEED (SElf-Evolving On-Policy Distillation), a self-evolving framework that converts completed on-policy trajectories into training-time hindsight skills and distills their behavioral effect back into the policy model. SEED first fine-tunes the policy to analyze completed trajectories and generate natural-language skills that capture reusable workflows, decisive observations, or failure-avoidance rules. During RL, the current policy both collects trajectories and serves as the analyzer that extracts hindsight skills from them. Policy updates therefore improve subsequent decision making and skill analysis together, allowing hindsight supervision to evolve with the policy. SEED then re-scores the sampled actions under ordinary and skill-augmented contexts, converting the skill-induced probability shift into a dense token-level on-policy distillation signal. This signal is jointly optimized with outcome-based RL, keeping the auxiliary supervision aligned with the current trajectory distribution. Extensive experiments on text-based and vision-based agentic tasks show that SEED consistently improves performance and sample efficiency, exhibiting robust generalization to unseen scenarios. Our code is available at https://github.com/jinyangwu/SEED.

↑