arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共享学习率并非选择性在策略蒸馏中的中性控制

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

Chencheng Zhu

arXiv 2609.22109首次发表:更新:

发表机构

UNSW Sydney(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示选择性在策略蒸馏中共享学习率并非中性控制,选择器与学习率存在纠缠,影响比较结论,建议报告分支×学习率矩阵。

AI 中文摘要

选择性在策略蒸馏仅在选择器评分最高的token位置训练学生模型,文献中在单一共享学习率下比较不同选择器——这一控制被认为具有中性。我们证明事实并非如此。在GSM8K上使用LoRA(Qwen2.5-1.5B学生模型,7B教师模型),跨越8倍学习率网格,密集监督统计上平坦(波动1.8个百分点,p=0.26),而每个选择性分支随学习率变化:随机5%子集为5.4个百分点,全变差选择器为6.7个百分点,可教学性选择器高达17.7个百分点。因此,密集与选择性之间的比较结论在学习率1e-4时为10.1个百分点,但在5e-5时为5.1个百分点——这一2.0倍差异由协议视为背景的参数决定——并且选择器之间六对成对显著性调用中有两对在相邻学习率之间翻转,而无任何排名反转。我们将此称为选择器-学习率纠缠,并将其追溯到选择本身而非步长:AdamW更新幅度跟踪学习率,差异在2.2%以内,尽管各分支间梯度范数差异达15.5倍。一项预注册的冻结评分消融实验(选择由初始学生模型评分;标准、预算和在策略rollout不变;每单元12个种子)表明,实时评分增加了3.79±1.69个百分点的学习率敏感性(p=0.035),而冻结分支仍显著纠缠(p=0.015):反馈循环加剧了该现象而非导致它。在文献实际使用的学习率(1e-6至1e-5)下进行全量微调时,该模式加剧:密集本身波动19.8个百分点,选择性分支波动49.5个百分点,结论从已发表工作点的不显著+3.6个百分点到高一个档次的+34个百分点(p=0.005)。在MATH-500上,LoRA下的学习率依赖性未复现,限定了该结果的范围,而选择性训练约10个百分点的成本则复现。我们建议报告分支×学习率矩阵,而非共享学习率列,作为选择器比较的前提条件。

英文摘要

Selective on-policy distillation trains a student only at the token positions a selector scores highest, and the literature compares selectors under a single shared learning rate--a control chosen to be neutral. We show it is not. Under LoRA on GSM8K (Qwen2.5-1.5B student, 7B teacher), across an 8x learning-rate grid, dense supervision is statistically flat (swing 1.8 pp, p=0.26) while every selective arm moves with the rate: 5.4 pp for a random 5% subset, 6.7 pp for a total-variation selector, up to 17.7 pp for a teachability selector. Consequently the dense-versus-selective verdict reads 10.1 pp at lr=1e-4 but 5.1 pp at 5e-5--a 2.0x difference decided by a parameter the protocol treats as scenery--and two of six pairwise significance calls between selectors flip between adjacent rates without any rank inversion. We call this selector-rate entanglement and trace it to selection itself rather than step size: AdamW update magnitudes track the rate to within 2.2% despite 15.5x gradient-norm differences across arms. A preregistered frozen-scoring ablation (selection scored by the initial student; criterion, budget, and on-policy rollouts unchanged; 12 seeds per cell) shows live scoring adds 3.79+/-1.69 pp of rate sensitivity (p=0.035) while the frozen arm remains significantly entangled (p=0.015): the feedback loop aggravates the phenomenon rather than causing it. Under full fine-tuning at the rates this literature actually uses (1e-6 to 1e-5) the pattern grows: dense itself swings 19.8 pp, the selective arm 49.5 pp, and the verdict ranges from a non-significant +3.6 pp at the published operating point to +34 pp (p=0.005) one notch hotter. On MATH-500 the rate dependence does not reproduce under LoRA, scoping that result, while the ~10 pp cost of selective training does. We prescribe reporting the arm x rate matrix, not a shared-rate column, as a precondition for selector comparisons.

Comments16 pages, 3 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑