在线蒸馏结合离线GRPO:训练紧凑的指令跟随重排序器
On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
浏览论文内容
中文总结 AI 辅助
本文提出结合离线GRPO与在线蒸馏的两阶段框架,训练出1B规模的紧凑指令跟随重排序器,在MAIR-11、MAIR-Full等基准上优于多种蒸馏方法及7B规模RL训练重排序器,实现了更优的质量效率权衡。
中文摘要 AI 辅助
紧凑的指令跟随重排序器在部署时颇具吸引力,但传统蒸馏流水线通常通过在固定示例集上离线模仿教师输出训练学生,将监督限制在教师观测到的排序空间内。本文从强化学习视角重新审视重排序器蒸馏,提出一种结合离线教师优化与在线学生蒸馏的两阶段框架:第一阶段,采用4B规模的教师重排序器,利用88K个指令跟随示例的LLM评判反馈,通过离线GRPO(Generative Reward Preference Optimization,生成式奖励偏好优化)进行强化;第二阶段,采用1B规模的紧凑学生模型,从自身策略中采样排序,并在这些排序上接收教师生成的软奖励,将学生探索与知识传递相结合。在分布偏移场景下,该方法的性能提升最为显著:在含11个子集、869个查询的MAIR-11原始评估集上,所提学生模型达到0.7670的nDCG@6,较离线列表式KD(Knowledge Distillation,知识蒸馏)高出4.6个百分点。与离线成对RankNet KD、在线GKD(Generative Knowledge Distillation,生成式知识蒸馏)的对比实验显示,无论是改变离线蒸馏目标,还是将教师分布匹配移至在线,均无法复现基于奖励的在线蒸馏在学生采样排序上的性能优势。该优势在含126个任务、9356个查询的MAIR-Full数据集上依然存在:在所评估的各类蒸馏变体中,所提方法取得最高的任务宏观点估计值,达到0.6808的nDCG@6与0.7865的MRR@6;在可比的MAIR-11评估集上,其性能优于两款已发布的7B规模经RL训练的重排序器,且相同的第二阶段训练流程可持续改进三种架构各异的替代学生骨干网络。在含9861个查询的验证基准上,生成的1B规模重排序器达到0.7624的nDCG@6,同时相比更大规模模型实现了更优的质量-效率权衡。
英文摘要
Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of examples, constraining supervision to the teacher's observed ranking space. We revisit reranker distillation through the lens of reinforcement learning. We propose a two-stage framework combining off-policy teacher optimization with on-policy student distillation. In Stage 1, a 4B teacher reranker is strengthened with off-policy GRPO using LLM-judge feedback on 88K instruction-following examples. In Stage 2, a compact 1B student samples rankings from its own policy and receives soft teacher-derived rewards on those rankings, coupling student exploration with knowledge transfer. Our strongest gains appear under distribution shift. On MAIR-11, the original 11-subset, 869-query evaluation, the proposed student reaches 0.7670 nDCG@6, outperforming offline listwise KD by +4.6 points. Controlled comparisons against offline pairwise RankNet KD and on-policy GKD show that neither changing the offline distillation objective nor moving teacher-distribution matching on-policy reproduces the performance of reward-based on-policy distillation over student-sampled rankings. The advantage persists on MAIR-Full: across all 126 tasks and 9,356 queries, the proposed method obtains the highest task-macro point estimates among the evaluated distillation variants, reaching 0.6808 nDCG@6 and 0.7865 MRR@6. It also exceeds two released 7B RL-trained rerankers on the comparable MAIR-11 evaluation, while the same Stage 2 training procedure consistently improves three architecturally distinct alternative student backbones. On the 9,861-query validation benchmark, the resulting 1B reranker achieves 0.7624 nDCG@6 while providing a favorable quality-efficiency tradeoff relative to larger alternatives.
发表机构
- SAP Labs(思爱普实验室)
机构由 AI 辅助整理,请以论文原文为准。