发表机构
Meta AI; UNC Chapel Hill(Meta AI; 北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Diffusion-GR2,通过转换微调、在线策略蒸馏和强化学习将自回归推理重排序器转换为块扩散模型,在保持精度的同时将推理吞吐量提升2.4-3.5倍。
AI 中文摘要
生成式推理重排序器通过在重新排序候选列表之前生成思维链来实现强大的推荐准确性,但推理速度较慢:自回归解码器每个推理令牌需要一个顺序前向传递,且推理轨迹远超其产生的排序。为降低这一成本,块扩散语言模型在少量去噪步骤中并行解码多个位置,速度显著提升,但简单地将自回归重排序器转换为块扩散模型会带来两个精度差距:(1) 结构差距:答案位置被并行去噪并独立评分,导致解码器产生无效排序(重复、遗漏或超出集合的标识符),而自回归通过从左到右掩码避免此问题;(2) 分布差距:在固定教师轨迹上微调转换后的模型相对于其推理时的自身解码是离策略的,留下残余精度差距。为在保持加速的同时弥合这两个差距,我们提出\textbf{Diffusion-GR2},一种将我们的自回归推理重排序器(GR2)转换为块扩散重排序器的方案。首先,转换微调使自回归初始化的扩散模型适应于独立将答案去噪为有效排列,无需外部约束解码器。其次,在线策略蒸馏利用来自自回归教师的密集逐令牌目标,在模型自身的解码轨迹上监督模型。最后,我们在在线策略蒸馏的在线策略策略之上应用针对重排序奖励的强化学习阶段。在Amazon Beauty上的实验表明,Diffusion-GR2恢复至与自回归重排序器接近的性能,而块并行解码在模型推理输出长度上将解码吞吐量提升2.4-3.5倍。消融实验显示,转换微调恢复了大部分转换差距,而在线策略蒸馏进一步将其缩小至自回归参考水平。
英文摘要
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces. To reduce this cost, block-diffusion language models decode many positions in parallel over a few denoising steps and are substantially faster, yet naively converting an AR re-ranker into one opens two accuracy gaps: (1) a structural gap: answer positions are denoised in parallel and scored independently, so the decoder emits invalid rankings (duplicated, dropped, or out-of-set identifiers) that AR avoids through left-to-right masking; and (2) a distributional gap: fine-tuning the converted model on fixed teacher trajectories is off-policy relative to its own decoding at inference, leaving a residual accuracy gap. To close both gaps while keeping the speedup, we propose \textbf{Diffusion-GR2}, a recipe that converts our AR reasoning re-ranker (GR2) into a block-diffusion re-ranker. First, conversion fine-tuning (CFT) adapts the AR-initialized diffusion model to denoise the answer into a valid permutation on its own, without an external constrained decoder. Next, on-policy distillation (OPD) then supervises the model on its own decoded trajectories with dense per-token targets from the AR teacher. Finally, we apply a reinforcement-learning (RL) stage against a re-ranking reward on top of OPD's on-policy policy. Experiments on Amazon Beauty demonstrate that Diffusion-GR2 recovers to near-parity with the AR re-ranker, while block-parallel decoding raises decode throughput by $2.4$--$3.5\times$ at the model's reasoning output length. Ablations show that CFT recovers most of the conversion gap, and that on-policy distillation further closes it to the AR reference.
CommentsWork in progress