arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11012cs.CL

EasyOPD:一种易于使用的大语言模型在线策略蒸馏框架

EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models

Jie Sun, Mao Zheng, Mingyang Song, Qiyong Zhong, Gengsheng Li, Zhepei Hong, Chang Wu, Pengfei Liu, Junfeng Fang, Xiang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对传统语言模型蒸馏问题,提出基于verl构建的EasyOPD框架,分离用户配置等,为三种OPD设置实例化方法,经多基准测试,其实现可通过同一后端执行,还发布了相关配置、文档及演示包。

中文摘要 AI 辅助

传统语言模型蒸馏通常依赖固定的教师生成数据,可能无法涵盖不断演变的学生策略遇到的状态。在线策略蒸馏(OPD)则对学生生成的轨迹收集教师或评估监督。但现有OPD方法在监督形式、分词器兼容性、教师访问和监督粒度上有很大差异,导致难以重现和扩展的碎片化实现。我们提出了EasyOPD,一个基于verl构建的在线策略蒸馏框架。EasyOPD分离了用户端配置、特定方法的监督逻辑和基于verl的执行。其方法模块通过扩展边界连接到共享后端进行损失构建、轨迹元数据、奖励处理、分词器对齐和教师端计算。我们为三种OPD设置实例化了代表性方法。实验表明这些实现可通过相同的基于verl的后端执行,同时保留其特定方法的目标和任务相关的性能配置文件。我们发布了带有可运行YAML配置、文档以及可安装演示包和视频的EasyOPD。

英文摘要

Conventional language-model distillation often relies on fixed teacher-generated data, which may not cover the states encountered by an evolving student policy. On-policy distillation (OPD) instead collects teacher or evaluator supervision on student-generated rollouts. However, existing OPD methods differ substantially in supervision form, tokenizer compatibility, teacher access, and supervision granularity, leading to fragmented implementations that are difficult to reproduce and extend. We present \textsc{EasyOPD}, an on-policy distillation framework built on verl, a distributed reinforcement-learning framework for large language models. \textsc{EasyOPD} separates user-side configuration, method-specific supervision logic, and verl-based execution. Its method modules connect to the shared backend through extension boundaries for loss construction, rollout metadata, reward processing, tokenizer alignment, and teacher-side computation. We instantiate representative methods for three OPD settings -- cross-tokenizer OPD, on-policy self-distillation, and step-wise OPD. Experiments on reasoning, code-generation, scientific-knowledge, and tool-use benchmarks show that these implementations can be executed through the same verl-based backend while retaining their method-specific objectives and task-dependent performance profiles. We release \textsc{EasyOPD} with runnable YAML configurations, documentation, and an installable demonstration package and video.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Tencent(腾讯)
  • Shanghai Innovation Institute(上海创新研究院)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑