发表机构
College of Elite Engineers, Nankai University; Zhongguancun Academy; Beijing Institute of Technology; Zhejiang University; Institute of Automation, Chinese Academy of Sciences; Harbin Institute of Technology; College of Software, Nankai University(南开大学精英工程师学院; 中关村学院; 北京理工大学; 浙江大学; 中国科学院自动化研究所; 哈尔滨工业大学; 南开大学软件学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型训练后处理问题,提出蒸馏强化学习方法,集成教师监督到RL目标,含反向重要性采样等三个组件,能有效转移知识,在家族内和跨家族蒸馏实验中性能大幅优于标准RL和OPD。
AI 中文摘要
大语言模型训练后处理对于提升推理、适应性和对齐性至关重要。现有方法主要有强化学习(RL)和策略蒸馏(OPD)两种范式。RL依赖粗粒度结果监督,存在信用分配困难和获取新知识能力有限的问题。OPD通过KL散度无条件匹配教师逻辑,导致类似教师提供新知识少,差异大的教师指导无效,限制其用于家族内蒸馏。我们提出蒸馏强化学习(Distilled RL),将教师监督集成到RL目标中,提供细粒度指导,选择性转移新知识并避免无条件模仿。它包含反向重要性采样、负样本重置和序列级几何归一化三个组件。通过案例研究表明其能有效转移知识,实验显示在家族内和跨家族蒸馏设置中,Distilled RL在pass@1和pass@k方面均大幅优于标准RL和OPD。
英文摘要
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into the RL objective to provide fine-grained guidance, selectively transfer new knowledge and avoid unconditional imitation. Distilled RL contains three components: reverse importance sampling with clipping, negative sample reset, and sequence-level geometric normalization. Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model. Extensive experiments across both within-family and cross-family distillation settings show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k. Our code is available at https://github.com/597358816/Distilled-RL.