arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12686cs.AIcs.CL

基于残差向量重建的长上下文召回,与上下文窗口大小无关

Residual Vector-based Reconstruction as Long-Context Recall Regardless of Context Window Size

MyungHoon Ryu, XinYu Piao, Jong-Kook Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种利用前馈层残差向量重建事实的长上下文召回方法,无需训练即可在保持近恒定GPU内存的同时,于两百万令牌上下文中成功回答单事实问题。

中文摘要 AI 辅助

大型语言模型(LLMs)处理长上下文,包括长文档和冗长对话,但面临随输入长度成比例增加的令牌级内存使用。尽管模型优化和有损提示压缩被广泛使用,这些方法仍无法解决超出预训练和大小受限上下文窗口的长上下文召回问题。本文提出一种长上下文召回方法,在上下文长度增加时保持近乎恒定的GPU内存使用,且无需额外训练。主要思想是利用LLM前馈层中的参数激活来重建事实,这些激活存储了表示源文档事实的残差向量。利用残差向量,LLM能够确定性地重建查询相关事实,而无需参考原始文档,从而保持高保真度并减少内存使用,无需微调权重。实验结果表明,所提方法能够在两百万令牌的故事上下文中回答单事实问题,而先前方法在此场景下失败。

英文摘要

Large language models (LLMs) process long contexts, including long documents and lengthy conversations, but face token-level memory usage that increases proportionally to input length. Although model optimization and lossy prompt compression are widely used, these methods still fail to solve the long-context recall problem beyond pretrained and size-constrained context windows. This paper proposes a long-context recall method that maintains near-constant GPU memory usage as context length increases, without additional training. The main idea is to reconstruct facts using parameter activations in the LLM's feed-forward layers, which store residual vectors representing facts from the source document. Utilizing residual vectors allows the LLM to deterministically reconstruct query relevant facts without referencing the original document, preserving high fidelity and reducing memory usage without fine-tuning weights. Experimental results show that the proposed method enables answering single-fact questions in two-million-token story contexts where previous methods fail.

发表机构

  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

↑