仅使用上下文示例的有效稠密检索
Effective Dense Retrieval using Only In-Context Examples
- University of Waterloo(滑铁卢大学)
- The University of Queensland(昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
提出免训练的RICE方法,利用上下文示例提示LLM生成高质量稠密表示,显著提升检索准确性,无需额外训练。
中文摘要 AI 辅助
将仅解码器的大型语言模型(LLM)转变为强大的稠密检索器通常需要某种形式的检索器训练。在本文中,我们探讨是否可以通过提示(prompting)仅利用少量上下文示例,使LLM生成有效的稠密检索表示。为此,我们提出了RICE(来自上下文示例的表示,Representations from In-Context Examples),一种简单的“免训练”方法,可从LLM中提取高质量的稠密表示。RICE通过让LLM基于提供查询和文档编码共享上下文的示例进行条件生成,从而生成表示。我们的结果表明,RICE嵌入能显著提高基于提示的LLM嵌入的准确性,使其成为一种无需训练的构建基于LLM的稠密检索器的简单方法。我们已在提供的URL上发布代码。
英文摘要
Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examples. To answer this, we introduce RICE (Representations from In-Context Examples), a simple "training-free" approach that extracts high-quality dense representations from LLMs. To do so, RICE conditions the LLM on examples that provide a shared context for query and document encoding. Our results demonstrate that RICE embeddings can substantially improve the accuracy of prompt-based LLM embeddings, establishing it as a simple method to build LLM-based dense retrievers that do not require training. We release our code at https://github.com/nourj98/RICE.