arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2404.05825cs.IRcs.AI

LLM增强检索:通过语言模型和文档级嵌入增强检索模型

LLM-Augmented Retrieval: Enhancing Retrieval Models Through Language Models and Doc-Level Embedding

  • Meta

机构由 AI 辅助整理,请以论文原文为准。

Mingrui Wu, Sheng Cao

更新

AI总结:

本文提出模型无关的LLM增强文档级嵌入框架,并改进负采样与损失函数等训练组件,显著提升Contriever、DRAGON和ColBERTv2等检索器,在LoTTE与BEIR上达到SOTA。

AI中文摘要:

近年来,与传统的稀疏或基于词袋的方法相比,基于嵌入的检索(或称稠密检索)已展现出最先进的结果。本文介绍了一种通过大型语言模型(LLM)增强、与模型无关的文档级嵌入框架。此外,本文还改进了检索模型训练过程中的一些重要组件,例如负采样、损失函数等。通过实现这一LLM增强的检索框架,我们已能够显著提升广泛使用的检索器模型(如Bi-encoders(Contriever、DRAGON)和后期交互模型(ColBERTv2))的有效性,从而在LoTTE数据集和BEIR数据集上取得最先进的结果。

英文摘要:

Recently embedding-based retrieval or dense retrieval have shown state of the art results, compared with traditional sparse or bag-of-words based approaches. This paper introduces a model-agnostic doc-level embedding framework through large language model (LLM) augmentation. In addition, it also improves some important components in the retrieval model training process, such as negative sampling, loss function, etc. By implementing this LLM-augmented retrieval framework, we have been able to significantly improve the effectiveness of widely-used retriever models such as Bi-encoders (Contriever, DRAGON) and late-interaction models (ColBERTv2), thereby achieving state-of-the-art results on LoTTE datasets and BEIR datasets.

↑