arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33409cs.CL

稠密并非足够:面向长程在线策略蒸馏的分层监督分配

Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation

  • Ant Group(蚂蚁集团)
  • Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团)
  • University of Science and Technology of China(中国科学技术大学)
  • Peking University(北京大学)
  • University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuhao Sun, Binrui Wu, Zhuoer Xu, Ming Wen, Haoxiang Xu, Bin Chen, Yan Lin, Qianzijing Zhang

AI总结:

针对长程在线策略蒸馏中均匀监督分配不佳的问题,提出分层框架LENS-OPD,通过定位、验证、精炼三个阶段,在多个智能体基准上显著提升任务性能。

AI中文摘要:

在线策略蒸馏(OPD)通过在学生自身轨迹上提供教师监督,将大型语言模型的能力迁移至较小的学生模型。然而,在长程智能体任务中,均匀的令牌级匹配可能导致监督分配不当:较大的局部差异未必能改善未来行为,而具有因果关联的指导可能超出当前学生的能力范围,或在缺乏特权输入时无法持续生效。我们将长程OPD形式化为分层监督分配问题,并主张有效指导位于未来效用与当前可学习性的交集处。关键在于,该交集随学生学习进程而演化。基于这一原则,我们提出LENS-OPD,一个由粗到细的框架,通过定位(Locate)、验证(Validate)和精炼(Refine)组织监督。定位阶段根据学生不断发展的能力调整轨迹暴露,并提出干预的候选决策。验证阶段测试该决策处的教师指导是否改善同一学生后续行为。精炼阶段将有益的指导行为内化至可部署策略中,并将令牌级监督集中于验证回合内决定性的师生冲突。这些阶段相互嵌套:每个更精细的分配以更粗的决策为条件,而非作为独立的重要性分数进行优化。在多个长程智能体基准及不同师生配置上的实验表明,LENS-OPD在任务性能上持续优于普通OPD及基于课程和选择的最强基线。我们的结果表明,有效的长程蒸馏需要在正确的深度、正确的决策和正确的令牌上进行教学。

英文摘要:

On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future behavior, while consequential guidance may be beyond the current student's reach or fail to persist without privileged input. We formulate long-horizon OPD as hierarchical supervision allocation and argue that productive guidance lies at the intersection of future utility and current learnability. Crucially, this intersection evolves as the student learns. Based on this principle, we propose LENS-OPD, a coarse-to-fine framework that organizes supervision through Locate, Validate, and Refine. Locate adapts trajectory exposure to the student's evolving competence and proposes a candidate decision for intervention. Validate tests whether teacher guidance at that decision improves the same student's subsequent behavior. Refine internalizes the beneficial guided behavior into the deployable policy and concentrates token-level supervision on decisive teacher-student conflicts within the validated turn. These stages are nested: each finer allocation is conditioned on the coarser decision, rather than being optimized as an independent importance score. Experiments across multiple long-horizon agent benchmarks and student-teacher configurations show that LENS-OPD consistently improves task performance over vanilla OPD and strong curriculum- and selection-based baselines. Our results suggest that effective long-horizon distillation requires teaching at the right depth, the right decision, and the right token.

↑