arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REIGN:采用集成引导网络的翻新嵌入用于高效上下文长度扩展

REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling

Devrim Çavuşoğlu, Emre Akbaş

arXiv 2608.29899首次发表:更新:

发表机构

Middle East Technical University; OBSS AI(中东科技大学; OBSS AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出REIGN模型,通过解耦处理与缓存嵌入降低训练成本,在多场景长文档检索中以更小参数规模达到与更大模型相当的性能

AI 中文摘要

长文档的密集检索成本高昂,标记级编码器的计算量随序列长度呈二次方增长,大多数长上下文嵌入模型仅通过架构变通方案或拉伸数十亿参数的大语言模型(LLM)才能达到32K标记。我们提出REIGN(Refurbished Embeddings with Integrated Guidance Networks,采用集成引导网络的翻新嵌入),这是一种对比训练的双编码器,它基于来自冻结引导网络(GN)的上下文块嵌入序列运行,而非原始标记。REIGN主要针对多块输入,用于文档到文档的检索;单块输入则直接使用GN。将标记级处理与文档级推理解耦,并将GN嵌入缓存到磁盘,与分块Transformer微调相比,每文档的训练成本降低了约四个数量级。我们还发布了一个合成长文档检索基准,用于长上下文长度下的对比训练与评估。在分布内的维基百科基准、LoCo分布外套件以及真实世界的专利检索案例研究中,REIGN在每种场景下都以更小的参数预算达到了与密集长上下文检索器相当的性能。配对显著性测试显示,在专利任务上,它与参数规模为其1.6-4.3倍的模型性能相当,在LoCo上,它与参数规模为其20倍的模型的nDCG@10差距不超过0.65。

英文摘要

Dense retrieval over long documents is expensive. Token-level encoders scale quadratically in sequence length, and most long-context embedding models reach 32K tokens only through architectural workarounds or by stretching billion-parameter LLMs. We propose REIGN (Refurbished Embeddings with Integrated Guidance Networks), a contrastively trained bi-encoder that operates on sequences of contextualised chunk embeddings from a frozen Guidance Network (GN) rather than on raw tokens. REIGN targets multi-chunk inputs, primarily for document-to-document retrieval; single-chunk inputs stay with the GN. Decoupling token-level processing from document-level reasoning, and caching the GN embeddings to disk, cuts per-document training cost by roughly four orders of magnitude relative to chunked Transformer fine-tuning. We also release a synthetic long-document retrieval benchmark for contrastive training and evaluation at long context lengths. Across an in-distribution Wikipedia benchmark, the LoCo out-of-distribution suite, and a real-world patent retrieval case study, REIGN matches dense long-context retrievers at smaller parameter budgets in each regime. A paired significance test puts it on par with models 1.6-4.3x larger on the patent task, and it stays within 0.65 nDCG@10 of a 20x-larger model on LoCo.

CommentsAccepted to Findings of EMNLP 2026. URL: https://devrimcavusoglu.github.io/reign/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑