arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从测试时缩放的视角理解在线策略蒸馏

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

Xinmu Ge, Zizhuo Zhang, Yu Huang, Jianing Zhu, Lin Yuan, Wanli Gu, Weichang Wu, Weiran Huang, Bo Han, Xiaolu Zhang, Jiangchao Yao

arXiv 2608.11829首次发表:更新:

发表机构

Hong Kong Baptist University; University of Texas at Austin; Shanghai Jiao Tong University; Shanghai Innovation Institute; Ant Group(香港浸会大学; 德克萨斯大学奥斯汀分校; 上海交通大学; 上海创新研究院; 蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究从测试时缩放视角分析OPD,发现其主要提升采样效率而非扩展推理能力边界,属于“虚假蒸馏”。

AI 中文摘要

在线策略蒸馏(OPD)已成为一种有前景的用于增强大语言模型(LLM)推理能力的后训练技术。人们普遍认为它能让学生模型从更强的教师模型中蒸馏知识,从而扩展超出OPD前基础模型的能力。本研究通过测试时缩放的视角,改变采样预算K并使用pass@K和avg@K评估性能来检验这一观点。具体而言,在多个OPD变体中,我们观察到经OPD训练的模型在所有采样预算下均保持更优的avg@K性能,而pass@K的优势随K增大逐渐转向OPD前的基础模型。这些结果表明OPD主要提升采样效率,而非持续扩展学生的推理能力边界。OPD训练过程中的pass@K动态进一步揭示,其以牺牲大K的能力边界为代价,逐步转向更强的小K性能。此外,以pass@1024为标准的问题级可解性分析揭示了一种不对称性:OPD导致更多此前可解的问题变为不可解,而非此前不可解的问题变为可解。综上,这些发现表明,从能力扩展视角看,OPD更像一种“虚假蒸馏”:其表观增益主要源于采样效率提升,而非从教师模型获取真正的新推理能力。

英文摘要

On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. Under the reverse KL objective, the idealized optimum of OPD aligns the student distribution with that of the teacher. When the teacher consistently outperforms the student, this naturally suggests that OPD should yield broad improvements over the pre-OPD student. However, do such improvements extend across the entire range of test-time sampling budgets? In this work, we revisit this expectation through the lens of test-time scaling by varying the sampling budget $K$ and evaluating performance with pass@$K$. Across multiple settings, we observe two distinct patterns: OPD can improve pass@$K$ at both small and large sampling budgets, but it can also improve small-budget performance while reducing large-budget pass@$K$. We show one condition that guarantees such a reversal and an idealized reverse KL counterexample where it occurs even when the teacher has higher accuracy on every problem. To choose between two candidate teachers at a target sampling budget, we propose the \textit{Teacher Advantage Score at $K$} (TAS@$K$), which can be computed before OPD training to predict which teacher will lead to a larger improvement in pass@$K$. Across three domains and thirteen benchmarks, the ordering predicted by TAS@$K$ agrees with the observed pass@$K$ improvements of the resulting OPD models in 83.6\% of experiments, providing a useful signal for teacher selection at the target pass@$K$.

Comments26 pages. Code and data: https://github.com/Geraldxm/opd-test-time-scaling; checkpoints: https://huggingface.co/collections/Geraldxm/opd-test-time-scaling-math-code-and-fact-checkpoints-6aba42275d3362d882cfc472

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑