AI 中文总结
本研究探讨在仅解码器Transformer中,共享全局KV缓存时保留层特定局部历史的价值,发现局部历史能降低困惑度,并提出了减少上层构建工作的后缀调度方法。
AI 中文摘要
仅解码器Transformer语言模型在生成过程中缓存键和值(KV)以复用过去的计算。跨层共享KV节省了存储,但降低了跨深度可用的表示多样性。我们研究了在共享全局KV的同时,局部内存应保留哪些内容,将历史内容与用于形成它的输入源分开。在126M参数和2K上下文的条件下,一项八种子研究发现,与当前令牌局部分支相比,使用局部历史可使保留测试困惑度降低约1.4%。容量、条目计数和训练计算控制支持历史内容的价值。在两种子比较中,当相邻层共享局部输入但保留独立投影时,该价值仍然存在;源共享还缩短了精确缓存构建依赖。与GQA和相邻层KV共享相比,相等的有界学习率搜索和新种子确认在更大缓存和更高长请求延迟下产生了更好的同源似然。在等令牌适应到8K后,与相邻层共享的排序仍然存在,但带有短上下文成本。八种子的外部书籍历史效应仍不确定,下游结果因任务而异。我们推导了一个足够的后缀调度,在精确算术中保留完整缓存的同时减少上层构建工作。
英文摘要
Decoder-only Transformer language models cache keys and values (KV) to reuse past computation during generation. Sharing KV across layers saves storage but reduces the diversity of representations available across depth. We study what local memory should retain alongside shared global KV, separating historical content from the input source used to form it. At 126M parameters and 2K context, an eight-seed study finds about 1.4% lower held-out test perplexity with local history than with a current-token local branch. Capacity, entry-count and training-compute controls support the value of historical content. In a two-seed comparison, this value persists when adjacent layers share local inputs while retaining independent projections; source sharing also shortens exact cache-construction dependencies. Against GQA and adjacent-layer KV sharing, equal bounded learning-rate searches and new-seed confirmation yield better same-source likelihood with larger caches and higher long-request latency. The ordering against adjacent-layer sharing persists after equal-token adaptation to 8K, with a short-context cost. The eight-seed external-book history effect remains uncertain, and downstream outcomes vary by task. We derive a sufficient suffix schedule that reduces upper-layer construction work while preserving the complete cache in exact arithmetic.
Comments41 pages, including supplementary material