发表机构
University of North Carolina at Chapel Hill; New York University(北卡罗来纳大学教堂山分校; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语言模型内存管理,将存储项保留决策转化为对项目是否被重用的估计问题,提出固定延迟平滑策略RMM。在受控设置中有优势,在第三方基准测试中表现一般,贡献了框架及测量与累积效果对比的映射。
AI 中文摘要
具有有限工作内存的语言模型必须反复决定保留哪些存储项。每种已部署的方法都在项目到达时做出决定,要么基于过去(StreamingLLM、H2O),要么基于对未来的猜测(SnapKV)。我们将此选择重新表述为关于隐藏信号(项目是否会被重用)的估计问题,将现有方法置于一个轴上,即提交延迟\(H\):在线过滤器和学习到的预测器在\(H = 0\)时提交,而Belady的离线最优值处于已知整个未来的位置。中间缺失的状态,即固定延迟平滑,会等待有限数量的步骤,观察正确的近期预测关注了哪些项目,然后才提交。这种测量证明了其效用,将Belady不可观察到的未来请求转化为我们从模型本身读取的内容。我们将其实例化为一种无需训练的策略RMM,它是H2O的严格推广,当测量均匀时恰好简化为H2O。在重用是内生且在时间上分离的受控设置中,证明的效用比累积注意力能更好地识别已使用的内存,并且小的有限内存表现得像大得多的内存。但在独立的第三方基准测试中,在NVIDIA的KVPress框架内与它自己的SnapKV、H2O和StreamingLLM实现进行比较时,优势大多消失:在单轮问答中,RMM与H2O相当,在流式多轮设置中输给了H2O和SnapKV。原因很简单:在自然文本上,模型对大多数令牌是正确的,所以按正确性加权注意力几乎不会改变它,并且证明的效用会退化为累积注意力,除非重用是明显且内生的,而标准基准测试并未体现这一点。我们的贡献是该框架以及关于测量何时胜过累积的真实映射,而不是新的技术水平。
英文摘要
A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag $H$: online filters and learned predictors commit at $H=0$, while Belady's offline optimum sits where the whole future is known. The missing regime in between, fixed-lag smoothing, waits a bounded number of steps, observes which items a correct near-future prediction attended to, and only then commits. This measurement, demonstrated utility, turns Belady's unobservable future request into something we read off the model itself. We instantiate it as a training-free policy, RMM, a strict generalization of H2O that reduces to it exactly when the measurement is uniform. In controlled settings where reuse is endogenous and separated in time, demonstrated utility identifies used memory far better than accumulated attention, and a small bounded memory behaves like a much larger one. But on independent third-party benchmarks, run inside NVIDIA's KVPress harness against its own SnapKV, H2O, and StreamingLLM implementations, the advantage mostly disappears: RMM is on par with H2O for single-turn question answering and loses to both H2O and SnapKV in a streaming multi-turn setting. The cause is simple: on natural text the model is correct about most tokens, so weighting attention by correctness barely changes it, and demonstrated utility collapses onto accumulated attention unless reuse is sharp and endogenous, which standard benchmarks do not exercise. Our contribution is the framework and an honest map of when measuring beats accumulating, not a new state of the art.
Comments8 pages, 3 figures, 3 tables