截断不良项,提升优良项:基于排名分类的BoN式知识蒸馏
Truncate Bad, Upweight Good: BoN-Style Distillation via Rank-Based Classification
- Technion–Israel Institute of Technology(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出TUP策略,通过截断低排名候选、重加权高排名项实现BoN式蒸馏,在离线对齐任务中表现与强基线相当。
AI中文摘要:
推理时选择方法(如Best-of-N,即从多个候选中选最优)通过采样一组候选并根据奖励模型选择排名最高的完成结果来提升生成质量。知识蒸馏旨在将该过程摊销为单一策略,用候选内排名替代原始奖励,学习对排名更高的完成结果赋予更高权重。然而,现有的基于排名的策略通常使用平滑的全支撑重加权,因此低排名的完成结果虽权重较低但仍保留在目标支撑中。尽管更尖锐的重加权会减少下尾质量,但也会增加对单一奖励模型在顶部做出的脆弱排名的依赖。我们提出TUP:一种截断不良项、提升优良项的策略,它从支撑中移除低排名的完成结果,并仅以可调锐度对保留的上尾进行重加权。TUP具有闭式的、与提示无关的归一化,可通过二元交叉熵完全离线训练,使用移位截断胜率作为软标签,以及蒸馏至参考的对数似然比作为对数几率。理论上,在某些假设下,我们证明对于任何未知的最优奖励,最佳单调排名重加权可通过下尾截断规则匹配,为移除下尾而非仅降低其权重提供形式支持。实验上,我们表明TUP与强大的离线对齐基线具有竞争力。
英文摘要:
Inference-time selection methods, such as Best-of-N, improve generation by sampling a pool of candidates and selecting the top-ranked completion according to a reward model. Distillation seeks to amortize this procedure into a single policy by replacing raw rewards with in-pool ranks and learning a policy that upweights higher-ranked completions. However, existing rank-based policies typically use smooth full-support reweighting, so low-ranked completions receive less mass but remain in the target support. Although a sharper reweighting reduces lower-tail mass, it also increases reliance on brittle ranking at the top made by a single reward model. We propose TUP: a Truncate-bad, Upweight-good Policy that removes low-ranked completions from the support and reweights only the retained upper tail with a tunable sharpness. TUP admits a closed-form, prompt-independent normalization and can be trained fully offline via binary cross-entropy, using shifted-truncated win-rates as soft labels and distilled-to-reference log-likelihood ratios as logits. Theoretically, under certain assumptions, we show that for any unknown oracle reward, the best monotone rank-reweighting can be matched by a lower-tail truncation rule, providing formal support for removing the lower tail rather than merely downweighting it. Empirically, we show that TUP is competitive with strong offline alignment baselines.