arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

隐式在线策略自蒸馏

Latent On-Policy Self-Distillation

Guibin Zhang, Jiayang Lyu, Ran Sun, Xinlei Yu, Haoyu Zhao, Qibing Ren, Shuicheng Yan

arXiv 2608.13040首次发表:更新:

发表机构

Shanghai Jiao Tong University; National University of Singapore(上海交通大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出隐式在线策略自蒸馏(LOPD),将教师的特权上下文改为端到端可学习,在智能体工具使用和代码生成任务上,以远低于对比方法的rollout预算实现了更优性能。

AI 中文摘要

让智能体从经验中学习并将其内化为策略,已成为自进化AI的核心问题。在线策略自蒸馏(OPSD)提供了一条有效途径,通过特权自我教师对智能体自身轨迹提供密集监督;然而,现有方法仍严重依赖设计者指定的特权人工制品(如答案、反馈、技能或轨迹),限制了持续自我改进所需的端到端可学习性和可扩展性。在本研究中,我们提出隐式在线策略自蒸馏(LOPD),该方法并非设计另一种带有新指定形式特权上下文的手动OPSD变体,而是使教师的特权上下文本身能从经验中进行端到端学习。技术上,LOPD检索相关经验并将其组合成连续隐式token,以调节自我教师;而学生从任务和交互历史中生成轨迹,并在每个访问的前缀处接收密集的token级监督。我们进一步引入特权间隔目标,以稳定和调节隐式上下文的学习。实验表明,LOPD具有(I)强大的性能,在智能体工具使用和代码生成任务上,优于RLVR以及代表性OPSD方法(包括OPSD、SDPO和Skill-SD);(II)高学习效率,以不到GRPO和Skill-SD 30%的rollout预算超越了这两种方法。消融研究进一步提供直接证据,表明使特权上下文可学习是实现这些增益的必要条件。总体而言,这些结果表明LOPD朝着更具可扩展性和自我导向的智能体进化范式迈出了一步。

英文摘要

Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervision on the student's own trajectories; however, existing methods still rely heavily on designer-specified privileged artifacts (e.g., answers, feedback, skills, or trajectories), limiting the end-to-end learnability and scalability required for continual self-improvement. In this work, we introduce Latent On-Policy Self-Distillation (LOPD), which, rather than proposing another hand-crafted OPSD variant with a newly prescribed form of privileged context, makes the teacher's privileged context itself learnable end-to-end from experience. Technically, LOPD retrieves relevant experiences and composes them into continuous latent tokens that condition a self-teacher, while the student generates trajectories from the task and interaction history and receives dense token-level supervision at every visited prefix. We further introduce a privileged-margin objective to stabilize and regulate the learning of latent context. Empirically, LOPD demonstrates (I) strong performance, outperforming RLVR and representative OPSD methods including OPSD, SDPO, and Skill-SD across both agentic tool use and code generation; and (II) high learning efficiency, surpassing GRPO and Skill-SD with less than 30% of their rollout budget. Ablation studies further provide direct evidence that making privileged context learnable is necessary for realizing these gains. Together, these results position LOPD as a step toward a more scalable and self-directed paradigm for agent evolution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑