通过知识蒸馏将语言模型转换为高效的交叉编码器用于RAG重排
Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking
浏览论文内容
中文总结 AI 辅助
研究针对交叉编码器推理成本高的问题,通过两阶段管道微调LLaMA 3为RAG重排器,先监督微调后4位量化,在特定领域问答基准上评估,相比交叉编码器基线有多项性能提升且降低推理开销。
中文摘要 AI 辅助
交叉编码器在检索增强生成(RAG)管道中实现了高重排准确率,但推理成本呈二次方增长,限制了实时部署。我们通过两阶段管道对LLaMA 3(8B)进行微调作为即插即用的重排器:首先通过Unsloth框架和LoRA适配器在自定义查询-文档相关性数据集上进行监督微调,然后进行4位量化以实现高效推理。结果模型取代了结合BM25和密集向量搜索的双检索器RAG管道中的交叉编码器。在特定领域问答基准上使用RAGAS框架评估,我们微调后的LLaMA 3重排器在答案相关性、上下文精度、答案相似度和答案正确性方面比交叉编码器基线分别提高了14%、16%、19%和21%,同时通过4位量化降低了推理开销。这些结果表明,经过指令微调的语言模型可以被改编为准确、高效的重排器,而无需传统交叉编码器的二次方复杂度。
英文摘要
Cross-encoders achieve high reranking accuracy in Retrieval-Augmented Generation (RAG) pipelines but impose quadratic inference costs that limit real-time deployment. We address this by fine-tuning LLaMA 3 (8B) as a drop-in reranker using a two-stage pipeline: supervised fine-tuning on a custom query-document relevance dataset via the Unsloth framework with LoRA adapters, followed by 4-bit quantization for efficient inference. The resulting model replaces the cross-encoder in a dual-retriever RAG pipeline combining BM25 and dense vector search. Evaluated on a domain-specific question-answering benchmark using the RAGAS framework, our fine-tuned LLaMA 3 reranker achieves gains of 14% in answer relevancy, 16% in context precision, 19% in answer similarity, and 21% in answer correctness over the cross-encoder baseline, while reducing inference overhead through 4-bit quantization. These results demonstrate that instruction-tuned LLMs can be adapted into accurate, efficient rerankers without the quadratic complexity of traditional cross-encoders.