arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LOCKS:用于高效长上下文解码的页面局部紧凑键摘要

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding

Junsung Hwang

arXiv 2607.24555首次发表:更新:

AI 中文总结

研究针对长上下文解码中KV缓存瓶颈问题,提出LOCKS方法,为页面提供局部紧凑键摘要,通过重建对数its、估计注意力质量等操作,仅关注顶部页面,在多种任务中表现出色,显著降低解码延迟。

AI 中文摘要

在长上下文环境中服务大型语言模型时,键值(KV)缓存成为瓶颈,每次解码步骤都要完整读取。注意力键虽全局高秩但局部低秩,共享低秩基会丢弃页面特定方向,而页面自身紧凑基能保留。LOCKS为每个页面提供其自身的频谱摘要(驻留,约为缓存大小的十分之一),重建页面内的对数its,通过对数求和指数估计每个页面的注意力质量,仅关注顶部页面;选择本身不读取候选键或值。在长文档问答(LongBench - v1)中,仅基于此摘要进行选择与完整缓存的差距在约一个点以内,在检索密集型RULER上能跟踪逐键读取的预言机直至最小预算,在长篇推理(AIME26,MATH - 500)中优势最大,此时基线选择器失效。在其2048令牌预算下,LOCKS在100K +上下文时与FullKV聚合质量匹配,同时关注约2%的令牌,与密集注意力相比,每令牌解码延迟减半(1M令牌时为2.0倍)。LOCKS作为未修改的vLLM的即插即用插件发布,批处理解码在完整的CUDA图中运行。

英文摘要

Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read at every decode step. We find that attention keys are approximately low-rank within pages. A single low-rank projection shared across pages can miss page-specific directions; fitting a basis to each page better identifies the pages receiving the most attention at comparable stored selector cost. LOCKS stores a rank-$r$ spectral summary per page, reconstructs its within-page logits, and selects pages by log-sum-exp mass without reading candidate keys or values. It stays within about a point of FullKV on LongBench-v1, tracks the read-every-key exact-LSE oracle on RULER down to the smallest budgets, and retains quality furthest under tight budgets on AIME26 and MATH-500. At a $2048$-token budget it matches FullKV aggregate quality beyond $100$K context while attending about $2\%$ of tokens. Across ranks $2$-$8$, summaries use $4$-$10\%$ of full-KV bytes. On GH200 with GPU-resident KV, LOCKS reduces complete decode-step time by $1.8\times$ at $512$K context. With full KV offloaded to Grace memory, it reaches $3.82$-$4.22\times$ the faster dense backend's aggregate throughput at $64$K-$256$K by serving larger batches.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑