arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

hLLM:用于生成式重排序的单遍解码

SPD: Single Pass Decoding for Generative Reranking

Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon

arXiv 2609.01807首次发表:更新:

发表机构

Meta Platforms, Inc.(元平台公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出hLLM,一种O(1)解码的格式专用策略,结合LoRA微调与教师排序蒸馏,使生成式重排序速度提升64倍,保持质量并连接组合优化,开辟实时排序新路径。

AI 中文摘要

大型语言模型(LLM)实现了最先进的生成式排序质量,但其生成的排序必须经过解码,而自回归解码会为每个生成的token花费一次顺序前向传播。我们注意到,排序器只需生成的token是对排序后项目进行命名的N个序数值,这种狭窄的排列结构输出格式允许采用比从左到右生成高效得多的解码策略。我们引入hLLM(匈牙利语言模型),一种针对格式优化的解码策略,它用O(1)次前向传播解码所有N个序数。hLLM通过轻量级自注意力头从LLM的预填充隐藏状态中读取N×K的项目-位置得分矩阵,然后通过匈牙利算法将序数解码为该矩阵的最优二分分配,从而通过构造生成有效排列而非通过修复。通过对训练信号和骨干适配的系统研究,我们表明基于LoRA的微调结合教师排序蒸馏实现了28毫秒的端到端推理,速度提升了64倍,同时保持与教师相当的排序质量。我们提供了完整的消融研究,分解了架构、训练信号和骨干适配的贡献。我们的框架将生成式排序与组合优化联系起来,为实时排序开辟了其他O(1)解码机制的路径。

英文摘要

Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the only tokens a ranker must emit are the $N$ ordinal values naming the items in ranked order, and that this narrow, permutation-structured output format admits decoding strategies which are much more efficient than left-to-right generation. We introduce SPD (Single Forward Pass), a format-specialized decoding strategy that decodes all $N$ ordinals in $O(1)$ forward passes. SPD reads an $N \times K$ item-position score matrix off the LLM's prefill hidden states with a lightweight self-attention head, then decodes the ordinals as the optimal bipartite assignment of that matrix via the Hungarian algorithm, yielding a valid permutation by construction rather than by repair. Through a systematic study of training signals and backbone adaptation, we show that LoRA-based fine-tuning combined with auto-regressive LLM ranking distillation reaches 28 ms end-to-end inference, a speed-up of 64x while maintaining ranking quality on par with the teacher. We provide a complete ablation decomposing the contributions of architecture, training signal, and backbone adaptation. Our framework connects generative ranking to combinatorial optimization, opening a path toward other $O(1)$-decode mechanisms for real-time ranking.

Comments10 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑