AI 中文总结
本研究系统探究同策略蒸馏中提示数量、来源与选择对迁移的影响,发现提示效用依赖教师-学生配对,随机采样为有效基线。
AI 中文摘要
同策略蒸馏(OPD)利用教师对其自身采样响应的反馈来训练学生,然而提示选择如何塑造跨教师-学生对的迁移仍鲜为人知。我们系统性地研究了跨RL和SFT续写对以及跨模型设置中的提示数量、来源和选择。我们发现OPD可以高度提示高效:少量提示即可接近大池性能,其中四个DAPO提示匹配了3,840个DeepMath提示所观察到的数学分数。然而,提示效用是关系性的而非内在的:仅更换教师即可逆转数学和代码提示的相对有效性。为了表征这些迁移差异,我们分析了跨提示支持和模型对的参数与功能变化。与教师的功能对齐在不同支持和目标任务间有所差异;在续写对中,与教师对齐的预测变化可以与弱参数对齐共存。在有效支持上的持续OPD可以在不利迁移后恢复性能。最后,针对性选择并不持续优于均匀随机采样,且单独表现不佳的源的过滤在三个配对支持抽取中未产生一致的增益。总体而言,我们的结果区分了提示效率与提示可互换性,并表明有效的数据选择取决于教师-学生对和目标能力,随机采样在所研究的设置中提供了有竞争力的基线。
英文摘要
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional changes across prompt supports and model pairs. Functional alignment with the teacher varies across supports and target tasks; in continuation pairs, teacher-aligned prediction changes can coexist with weak parameter alignment. Continued OPD on effective supports can restore performance after unfavorable transfer. Finally, targeted selection does not consistently outperform uniform random sampling, and filtering out a source that performs poorly alone yields no consistent gain across three paired support draws. Overall, our results distinguish prompt efficiency from prompt interchangeability and show that effective data choice depends on the teacher-student pair and target capability, with random sampling providing a competitive baseline in the studied settings.
Comments39 pages. Code: https://github.com/wyy-1112/dissecting-opd