arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35505cs.LG

OPD 的强化学习视角:用于样本高效 LLM 推理的最小二乘策略蒸馏

An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

  • University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
  • Brigham Young University(杨百翰大学)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang

AI总结:

本文从强化学习视角研究在线策略蒸馏,提出最小二乘策略蒸馏(LSPD)框架,通过乐观探索与离线数据复用提升策略多样性和样本效率,在数学推理基准上平均提升+1.59分,并实现显著更优的Pass@k性能。

AI中文摘要:

我们通过强化学习的视角研究在线策略蒸馏(OPD),建立了 OPD 中的反向 KL 目标与 KL 正则化策略优化之间的联系。基于这一联系,我们引入了最小二乘策略蒸馏(LSPD),这是一个受强化学习启发的框架,将基于值的强化学习中的乐观探索和离线策略数据复用引入策略蒸馏。LSPD 通过探索保持策略多样性,同时通过反复学习先前收集的轨迹来提高 rollout 效率。我们的理论分析将 LSPD 与乐观基于值的学习联系起来,并表明其理想化公式在在线探索下实现了尖锐的 $\tilde{\mathcal O}(\log K)$ 遗憾界。在实证方面,LSPD 在六个数学推理基准和多种师生设置中持续优于现有蒸馏基线,在 Avg@16 中平均提升 +1.59 分。值得注意的是,通过高达 k=64 的 Pass@k 评估,我们发现 LSPD 通过随着 k 的增长实现更强的性能而更好地保持了策略多样性。其完全离线策略变体仅使用前 25% 的 rollout 批次即可达到与普通 OPD 相当的性能。总之,这些结果为 OPD 提供了强化学习视角,既提供了原则性解释,也为更有效和 rollout 高效的语言模型蒸馏提供了实用途径。

英文摘要:

We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp $\tilde{\mathcal O}(\log K)$ regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.

补充信息

↑