arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15456cs.LGcs.CL

循环潜在注意力:用于循环变换器的跨循环键值压缩

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers

James O' Neill, Fergal Reid

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对循环变换器解码时键值缓存占用大的问题,提出循环潜在注意力(LLA)编解码器,通过存储紧凑潜在表示按需重建键值向量,经奇异值分解初始化和蒸馏细化,在缓存压缩上效果显著,提升了模型性能。

中文摘要 AI 辅助

循环的、权重绑定的变换器通过重用块来减少参数,但解码仍为每个循环步骤存储单独的键/值缓存。我们表明这种循环索引缓存具有高度结构化。对于固定令牌、层和头,键/值向量在循环中遵循短的低秩轨迹,而头和层轴则更平坦。我们引入循环潜在注意力(LLA),一种训练后缓存编解码器,它存储紧凑的键和值潜在表示,仅在注意力读取时重建特定循环的键/值向量。默认的 per-head 编解码器压缩循环,而 LLA-2D 在极端压缩情况下还将头折叠到一个潜在表示中。编解码器从教师激活的奇异值分解初始化,并用 KL 和注意力输出蒸馏进行细化。在匹配的缓存预算下,per-head LLA 优于头轴 MLA、跨层共享、键值量化和最终循环重用,表明循环缓存是低秩的,但不能安全地折叠到单个状态。相同的轴优势在 Ouro-2.6B-Thinking 上成立,并转移到 Huginn-3.5B 上,在解码器独立评估中,奇异值分解编解码器在 32 倍压缩下仍接近无损。缓存减少是精确的。在一个 H200 上,潜在存储路径将 Ouro-1.4B 在 4k 上下文下的测量批量容量从 32 提高到 768 个序列,压缩率为 21.3 倍。对于长数学展开,对学生生成前缀的策略内细化将 MATH-500 在 4 倍压缩下从 0.43 提高到 0.66,并减少无答案生成。

英文摘要

Looped, weight-tied Transformers reduce parameters by reusing a single block, but decoding still stores a separate K/V cache for every recurrence step. We show that this loop-indexed cache is highly structured. For a fixed token, layer and head, K/V vectors trace a short low-rank trajectory across loops, while the head and layer axes remain much flatter. We introduce Looped Latent Attention (\lla{}), a post-training cache codec that stores compact K and V latents and reconstructs loop-specific K/V vectors only when attention reads them. The default per-head codec compresses recurrence, while \lla{}-2D also folds heads into one latent for the extreme-compression regime. The codec is initialized from the SVD of teacher activations and refined with logit and attention-output distillation. At matched cache budget, per-head \lla{} outperforms head-axis MLA, cross-layer sharing, KV quantization and final-loop reuse, showing that the recurrent cache is low-rank but not safely collapsible to a single state. The same axis advantage holds on Ouro-2.6B-Thinking and transfers to Huginn-3.5B, where an SVD codec remains near-lossless to $32\times$ compression in decoder-independent evaluation. The cache reduction is exact. On one H200, the latent-store path increases measured Ouro-1.4B batch capacity at 4k context from 32 to 768 sequences at $21.3\times$ compression. Lastly, for long reasoning rollouts such as in MATH-500, on-policy refinement on student-generated prefixes raises accuracy at $4\times$ compression from 0.43 to 0.66 and reduces no-answer generations when compared to token-level off-policy distillation.

发表机构

  • Fin AI Research(金融人工智能研究)

机构由 AI 辅助整理,请以论文原文为准。

↑