发表机构
NVIDIA; Technion(英伟达; 以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出UNREAL框架,利用冻结LLM内部表示统一语料库检索与长上下文证据选择,仅增少量参数,显著提升检索召回率与长上下文任务性能。
AI 中文摘要
长上下文推理与检索增强生成(RAG)在证据选择上处理着规模迥异的场景,从单个长提示到整个语料库。我们探究是否有一种模型内部的机制能够跨越这一范围进行证据选择。为此,我们提出了统一检索与长上下文的单一模型(UNREAL),这是一种模型原生的证据选择框架,旨在同时覆盖语料库检索与长上下文推理。UNREAL对文本块进行编码,并直接从冻结的大语言模型(LLM)的内部表示中推导出检索查询。它仅增加不到50万个可训练参数,且保持骨干网络不变。在一个包含30亿词元、2100万文本块的维基百科索引上,所有四种稠密型和混合型UNREAL骨干模型均超越了最先进的检索器-重排序器系统。最佳模型在HotpotQA上将召回率从49.1%提升至73.2%,在2WikiMultiHopQA上从31.7%提升至60.1%,在MuSiQue上从8.8%提升至14.4%。应用于长上下文任务时,相同的选择机制在生成前移除干扰项,将NoLiMa在其最大上下文长度128K词元下的准确率从1.0%提升至24.83%,并将LV-Eval在256K下的F1分数从49.97%提升至54.66%。此外,与全上下文推理相比,UNREAL从约32K词元起减少了FLOPs和首词元时间,且随着上下文增长,收益更大。这些结果共同确立了模型内部的证据选择作为语料库检索和证据稀疏长上下文推理的共同基础。
英文摘要
Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.