arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11933cs.CLcs.IRcs.LG

通过知识蒸馏将语言模型转换为高效的交叉编码器用于RAG重排

Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking

Shreeya Dasa Lakshminath, Shubhan S

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对交叉编码器推理成本高的问题,通过两阶段管道微调LLaMA 3为RAG重排器,先监督微调后4位量化,在特定领域问答基准上评估,相比交叉编码器基线有多项性能提升且降低推理开销。

中文摘要 AI 辅助

交叉编码器在检索增强生成(RAG)管道中实现了高重排准确率,但推理成本呈二次方增长,限制了实时部署。我们通过两阶段管道对LLaMA 3(8B)进行微调作为即插即用的重排器:首先通过Unsloth框架和LoRA适配器在自定义查询-文档相关性数据集上进行监督微调,然后进行4位量化以实现高效推理。结果模型取代了结合BM25和密集向量搜索的双检索器RAG管道中的交叉编码器。在特定领域问答基准上使用RAGAS框架评估,我们微调后的LLaMA 3重排器在答案相关性、上下文精度、答案相似度和答案正确性方面比交叉编码器基线分别提高了14%、16%、19%和21%,同时通过4位量化降低了推理开销。这些结果表明,经过指令微调的语言模型可以被改编为准确、高效的重排器,而无需传统交叉编码器的二次方复杂度。

英文摘要

Cross-encoders achieve high reranking accuracy in Retrieval-Augmented Generation (RAG) pipelines but impose quadratic inference costs that limit real-time deployment. We address this by fine-tuning LLaMA 3 (8B) as a drop-in reranker using a two-stage pipeline: supervised fine-tuning on a custom query-document relevance dataset via the Unsloth framework with LoRA adapters, followed by 4-bit quantization for efficient inference. The resulting model replaces the cross-encoder in a dual-retriever RAG pipeline combining BM25 and dense vector search. Evaluated on a domain-specific question-answering benchmark using the RAGAS framework, our fine-tuned LLaMA 3 reranker achieves gains of 14% in answer relevancy, 16% in context precision, 19% in answer similarity, and 21% in answer correctness over the cross-encoder baseline, while reducing inference overhead through 4-bit quantization. These results demonstrate that instruction-tuned LLMs can be adapted into accurate, efficient rerankers without the quadratic complexity of traditional cross-encoders.

补充信息

↑