AI 中文总结
本文提出SMem架构,通过块局部编码和交叉注意力实现上下文缓存精确复用与删除,在长上下文检索和编辑效率上显著优于现有Transformer,同时保持困惑度可比。
AI 中文摘要
Transformer的KV缓存将每个token的表示与其整个前缀纠缠在一起:一个段落一旦编码,就不能在不同的前缀下复用,也不能在不重新计算其后所有内容的情况下被删除,因此精确的缓存复用仅限于共享前缀。我们提出SMem,一种其上下文表示通过构造即为缓存的架构。块局部编码器将每个块独立于其他块映射到内存行,读取器通过交叉注意力对其并集进行条件生成。对于每个参数设置,内存精确地在固定块索引处组合,删除一个块是对b-token块的精确O(b)更新,且内存状态与编辑路径无关。在训练上下文的4倍长度下,在共享配方下,SMem能检索到超出任何训练长度窗口的植入针(在31和63块距离处精确匹配0.14-0.28),而学习位置、RoPE和Block-Attention风格的Transformer得分均不超过0.02。完全缓存的上下文通过仅计算一个块即可服务,耗时近恒定3.1-6.2毫秒,而冷预填充随上下文增长;批量解码存储的KV行减少34-38%,在带宽受限时运行速度快1.4-1.7倍;删除在512块时比后缀重计算快8.5倍,在4096块时快452倍(训练长度的32-256倍,探测成本模型而非服务场景)。代价是与具有相同位置方案的参数匹配Transformer相比,困惑度差距为-4.7%至+2.8%(负值有利于SMem),在FineWeb-Edu上从160M到1.5B,跨越两种配方和学习率搜索。SMem还能与RoPE组合:在160M和410M规模下,组合模型匹配或领先于匹配的Transformer,并缩小了SMem与RoPE Transformer差距的29-59%。因此,去除前缀纠缠使困惑度保持可比,同时使缓存精确可组合和可编辑。
英文摘要
The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose context representation is a cache by construction. A block-local encoder maps each block to memory rows independently of other blocks, and a reader conditions generation on their union through cross-attention. For every parameter setting, memory composes exactly at fixed block indices, deleting a block is an exact $O(b)$ update for $b$-token blocks, and the memory state is independent of the edit path. At $4\times$ the training context, under the shared recipe, SMem retrieves planted needles beyond any trained-length window (exact match 0.14-0.28 at distances of 31 and 63 blocks), where learned-position, RoPE, and Block-Attention-style transformers all score at most 0.02. A fully cached context is served by computing one block alone at a near-constant 3.1-6.2 ms, whereas cold prefill grows with context; batched decode stores 34-38% fewer KV rows and runs 1.4-1.7$\times$ faster when bandwidth-bound; and deletion beats suffix recomputation by 8.5$\times$ at 512 blocks and 452$\times$ at 4096 blocks (32-256$\times$ the trained length, probing the cost model rather than a served regime). The cost is a perplexity gap of -4.7% to +2.8% (negative favors SMem) against a parameter-matched transformer with the same positional scheme, at 160M-1.5B on FineWeb-Edu across two recipes and a learning-rate search. SMem also composes with RoPE: at 160M and 410M the composite matches or leads the matched transformer and closes 29-59% of SMem's gap to a RoPE transformer. Dropping prefix entanglement thus keeps perplexity comparable while making the cache exactly composable and editable.