发表机构
LUMIA Lab, Shanghai Jiao Tong University; ByteDance; Shanghai Jiao Tong University(上海交通大学生涯智能计算实验室; 字节跳动; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM逐点重排序的文档侧冗余计算问题,提出DoPR压缩文档前缀框架,实现内存缩减与延迟加速,同时保留重排序性能。
AI 中文摘要
大语言模型(LLM)是有效的重排序器,但逐点重排序会在不同查询间重复处理同一文档,造成大量文档侧冗余计算。我们提出DoPR,一种将离线文档处理与在线重排序解耦的压缩文档前缀框架。DoPR首先选择与查询无关的文档表示,将其转换为压缩文档前缀状态,该状态可离线预计算并在文档被检索时复用。在线重排序期间,模型仅处理查询和得分token来对每个查询-文档对打分,文档信息由存储的前缀状态提供。该设计通过文档侧压缩和跨查询前缀状态复用降低在线成本。在TREC DL、BEIR和BRIGHT数据集上,使用0.6B至8B的Qwen3模型开展的实验显示,DoPR实现了最高8.0倍的在线文档侧内存缩减和最高8.04倍的延迟加速,同时保留了匹配全文档重排序器97.1%-99.5%的平均NDCG@10。
英文摘要
Large language models (LLMs) are effective rerankers, but pointwise reranking repeatedly processes the same document across different queries, causing substantial redundant document-side computation. We propose \textbf{DoPR}, a compressed document prefix framework that decouples offline document processing from online reranking. DoPR first selects query-independent document representations and converts them into compressed document prefix states, which are precomputed offline and reused whenever the document is retrieved. During online reranking, the model scores each query-document pair by processing only the query and scoring token, with document information supplied by the stored prefix states. This design reduces online cost through both document-side compression and cross-query prefix-state reuse. Experiments on TREC DL, BEIR, and BRIGHT with Qwen3 models from $0.6$B to $8$B show that DoPR achieves up to 8.0$\times$ online document-side memory reduction and up to 8.04$\times$ latency speedup, while retaining \textbf{97.1\%-99.5\%} of the average NDCG@10 of matched full-document rerankers.
Comments15 pages, 6 figures