arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38987cs.LG

更小的模型,更好的拒绝样本:偏好蒸馏的规模效应

Smaller Models, Better Rejects: Preference Distillation Scaling

Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov… 展开作者

Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov, Kayhan Behdin, Zhengze Zhou, Ran He, Alborz Geramifard, Rohit Jain, Zhe Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过理论和实验证明,偏好蒸馏中较小冻结模型生成的拒绝样本比学生自生成样本更有效且成本更低,并提出混合来源、重分配及低似然选择等干预策略以提升训练效果。

中文摘要 AI 辅助

偏好蒸馏通常将教师模型的回答视为首选,而将学生模型自身的回答视为被拒绝的样本。这一做法假设了自生成的失败样本是最具信息量的负样本,并且拒绝样本必须来自至少与学生模型同等规模的模型,从而使得大规模生成成本高昂。我们发现这两个假设均不成立:在从7B到72B的学生模型中,较小的冻结模型生成的拒绝样本所需推理计算更少,但在代码生成和数学推理任务上,无论是在序列级知识蒸馏之前还是之后,这些样本训练出的学生模型都强于使用自生成拒绝样本训练出的模型。为了解释这一结果,我们在线性化特征模型中推导了直接偏好优化的有限时域效用界。该界刻画了有利的拒绝样本分布,并激发了三种干预措施。首先,将来自较小模型和学生规模模型的拒绝样本混合使用,随着较小模型占比的增加,性能会得到提升。其次,将拒绝样本重新分配给其他提示并打乱其代码标记,其效果仍优于长度匹配的乱码,这表明任务结构对拒绝样本的效用有所贡献。第三,在参考策略下选择似然较低的候选样本,当似然较高的候选样本提供的对比信息较少时,能够改善净迁移效果。对于每一种来源,低似然选择的表现都优于高似然选择。这些结果表明,有效的拒绝样本在保持任务结构的同时,限制了与参考策略的耦合,而较小的冻结模型能够以低成本提供这些样本。

英文摘要

Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To explain this result, we derive a finite-horizon utility bound for Direct Preference Optimization in a linearized feature model. The bound characterizes favorable reject distributions and motivates three interventions. First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, reassigning rejects to other prompts and shuffling their code tokens still outperform length-matched gibberish, showing that task structure contributes to reject utility. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source. These results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.

发表机构

  • LinkedIn(领英)
  • University of California, Davis(加州大学戴维斯分校)
  • Arizona State University(亚利桑那州立大学)
  • Clemson University(克莱姆森大学)
  • Pennsylvania State University(宾夕法尼亚州立大学)
  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

↑