发表机构
UC Berkeley; UT Austin(加州大学伯克利分校; 得克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文首次系统研究百万Token语料库上的上下文检索,提出0.6B参数的BlockSearch模型,通过注意力稀释分析引入长度感知调整,在MS MARCO和NQ上匹配稠密检索,在LIMIT上得分高出3倍。
AI 中文摘要
语言模型(LMs)为基于向量的检索提供了一种有趣的替代方案:在上下文语料库上进行条件化并直接生成相关答案。然而,先前的工作主要集中于专有系统或较小规模的重新排序任务,而语料库规模的上下文检索在很大程度上尚未被探索。在这项工作中,我们首次对实际检索器所需的两个尺度上的上下文检索进行了系统研究:百万Token语料库和远超训练时大小的长度泛化。我们首先引入BlockSearch,一个0.6B的LM检索器,其架构和训练改进优于先前的LM基线,并在训练范围外长度泛化高达10倍。然而,在更极端的推断下,检索仍然崩溃。我们将这种失败归因于注意力稀释效应:随着语料库增长,不相关文档主导了softmax分母,即使黄金文档的预softmax分数保持高位,其归一化质量也会降低。基于这一分析,我们引入了对注意力softmax和文档级稀疏注意力的长度感知调整。通过这些修改,在百万Token规模下,我们的模型在广泛研究的基准(如MS MARCO和NQ)上匹配稠密检索,同时尽管比并发模型MSA小7倍,但性能更优。此外,在需要完全不同相似性概念的任务(如LIMIT)上,它显著优于稠密检索,得分高出3倍。综合来看,我们的结果将上下文检索定位为经典检索的有前景替代方案,同时强调在极端上下文增长下的注意力控制是一个新的挑战。
英文摘要
Language models (LMs) raise an intriguing alternative to vector-based retrieval: conditioning on an in-context corpus and directly generating a relevant answer. However, prior work has largely focused on proprietary systems or the smaller-scale reranking task, leaving corpus-scale in-context retrieval largely unexplored. In this work, we present the first systematic study of in-context retrieval on two scales practical retrievers demand: million-token corpora and length-generalization far beyond training-time sizes. We first introduce BLOCKSEARCH, an 0.6B LCLM retriever whose architectural and training modifications improve over prior LM baselines and length-generalize up to 10x beyond its training regime. Nevertheless, its retrieval still collapses under more extreme extrapolation. We trace this failure to an attention dilution effect: as the corpus grows, irrelevant documents dominate the softmax denominator and the normalized mass on the gold document collapses. Individual attention heads continue to locate the gold document more reliably than the model decodes it, even at million-token scale, though this signal also weakens as the corpus grows. Motivated by this analysis, we introduce length-aware adjustments to the attention softmax and document-level sparse attention, improving retrieval at million-token scale to performance comparable to a same-backbone dense retriever. On the lexical LIMIT benchmark, these gains also transfer out of distribution: at a million-token corpus, Recall@1 reaches 13.5% versus 2.9% for the dense baseline. Together, our results position in-context retrieval a promising alternative to classical retrieval while emphasizing attention control under extreme context growth as a new challenge.
Commentsaccepted to NeurIPS 2026