arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环网络中的动态压缩

Dynamic Compression in Recurrent Networks

Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal

arXiv 2608.17896首次发表:更新:

发表机构

Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出动态压缩方法,允许循环模型选择性重访过去标记以修正固定状态,在少样本函数复用任务中,其所需循环状态更小、扩展性更优,实现了计算与内存的有效权衡。

AI 中文摘要

循环模型通过将历史信息压缩为固定大小的状态来高效处理长上下文,但现代架构通常在序列上进行单次因果传递,因此每个输入必须在模型知晓其后续用途前完成压缩,迫使有限的状态在未来可能的需求间做出妥协。我们引入动态压缩,它允许循环模型有选择性地重新访问过去的标记,并通过额外的循环更新修正其固定大小的状态。模型无需以均匀的高保真度在循环状态中保留历史的每个部分,因为低保真度信息在相关时可从保留的原始序列中重新获取。我们在受控场景中研究该机制:模型首先在上下文内学习多个函数,随后在同一序列中遇到一系列少样本任务,每个任务要求模型识别并复用其中一个函数。单次传递模型必须以足够的保真度保留所有函数以应对任何未来任务,而选择性重新扫描允许模型仅重新访问并修正当前所需的函数。我们发现,动态压缩大幅降低了准确复用所需的循环状态大小,且随着存储函数数量的增加,其扩展性更优。这些结果表明存在一种计算-内存权衡:循环模型可投入更多计算重新访问历史,以更有效地利用固定大小的状态。

英文摘要

Recurrent models process long contexts efficiently by compressing their history into a fixed-size state, but modern architectures typically do so in a single causal pass over the sequence. Each input must therefore be compressed before the model knows how it will later be used, forcing a limited state to compromise across possible future demands. We introduce dynamic compression, which allows a recurrent model to selectively revisit past tokens and revise its fixed-size state through additional recurrent updates. The model need not preserve every part of the history at uniformly high fidelity in its recurrent state, because lower-fidelity information can be revisited from the retained raw sequence when it becomes relevant. We study this in a controlled setting where the model first learns multiple functions in-context and, later in the same sequence, encounters a series of few-shot tasks that each require it to identify and reuse one of those functions. A single-pass model must preserve every function at sufficient fidelity for any future task, whereas selective re-scanning allows the model to revisit and refine only the function currently needed. We find that dynamic compression substantially reduces the recurrent state required for accurate reuse and scales more favorably as the number of stored functions grows. These results demonstrate a computation--memory tradeoff in which recurrent models can spend more computation revisiting their history to make more effective use of a fixed-size state.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑