arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07282cs.CL

将流稳定性与语言模型中的长期记忆分离

Separating Stream Stability from Long-Term Recall in Language Models

  • School of Computer and Electronic Information, Guangxi University(广西大学计算机与电子信息学院)
  • School of Information Science and Engineering, Chongqing Jiaotong University(重庆交通大学信息科学与工程学院)
  • School of Computer Science and Technology, Guangdong University of Technology(广东工业大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Peipei Cao, Xin Zhang, Jie Tang, Xiao Li, Siying Li, Qing Pei

AI总结:

本研究提出将流式语言模型的稳定性与长期记忆分离,通过定义稳定性、访问和效用三个视界及ThreeH评估框架,证明注意力汇聚仅保证无限稳定性但非长期记忆,循环和检索可扩展语义视界。

AI中文摘要:

流式语言模型的方法常常与长上下文和记忆系统一起讨论,尽管它们解决的是不同的问题。注意力汇聚(attention sink)可以在无限长的流上稳定自回归生成,而模型仍然无法使用已离开其近期令牌缓存的内容。我们认为,这种区别应在系统声明和评估中明确。我们引入了三个视界:稳定性视界(stability horizon),在此范围内预测行为保持良好;访问视界(access horizon),在此范围内过去的内容仍能因果影响输出;以及效用视界(utility horizon),在此范围内任务保持可接受的性能。我们建设性地表明,稳定性视界可以是无限的,而访问视界和效用视界是有限的。然后,我们提出了ThreeH,一个在共同状态和计算预算下测量所有三个视界的评估契约。将该框架应用于注意力汇聚流式,阐明了其优势:恒定内存、稳定生成,而不将锚定令牌视为语义记忆。该框架揭示了缓存策略、循环状态、检索和外部记忆的作用。在128K令牌流、延迟绑定回忆和延迟决策上的实验表明,注意力汇聚保留了局部建模,但不保留活动缓存之外的内容;循环和检索状态扩展了语义视界。

英文摘要:

Methods for streaming language models are often discussed alongside long-context and memory systems, although they solve different problems. An attention sink can stabilize autoregressive generation over an indefinitely long stream while the model remains unable to use content that has left its recent-token cache. We argue that this distinction should be explicit in system claims and evaluation. We introduce three horizons: the stability horizon, over which predictive behavior remains well behaved; the access horizon, over which past content can still causally affect the output; and the utility horizon, over which a task retains acceptable performance. We show constructively that the stability horizon can be infinite while the access and utility horizons are finite. We then propose ThreeH, an evaluation contract that measures all three horizons under a common state and compute budget. Applying the framework to attention-sink streaming clarifies its strength, constant-memory, stable generation, without treating anchor tokens as semantic memory. The framework exposes roles for cache policies, recurrent state, retrieval, and external memory. Experiments on 128K-token streams, delayed binding recall, and delayed decisions show that attention sinks preserve local modeling but not content beyond the active cache; recurrent and retrieval state extend the semantic horizon.

↑