AI 中文总结
针对长上下文LLMs自注意力的解码效率瓶颈,提出无训练的BinaryPC方法,通过数据二元主成分构建哈希码,在保持精度的同时将解码吞吐量提升3.56倍。
AI 中文摘要
长上下文大语言模型(LLMs)正越来越多地被部署到实际应用中,但自注意力机制仍是主要的效率瓶颈——尤其是在解码阶段,因为需要重复处理不断增长的键值(KV)缓存。现有的稀疏注意力机制通过关注更少的KV对来减少计算,但通常会出现显著的精度下降、需要额外训练或依赖昂贵的哈希方法。在本研究中,我们提出了BinaryPC,一种用于长上下文LLMs的无训练、数据感知的基于哈希的稀疏注意力机制。BinaryPC通过计算数据的二元主成分来构建紧凑的二元哈希码和对应的哈希函数。与基于数据无关随机投影的局部敏感哈希(LSH)或学习型非线性哈希方法不同,BinaryPC构建的二元码能明确保留数据的结构信息,且无需基于梯度的训练。在多个模型系列和长上下文基准上的综合实验表明,BinaryPC在相对于全注意力机制保持精度的同时,在稀疏和基于哈希的基线方法中实现了更优性能。在现代GPU上,BinaryPC相比FlashAttention内核将端到端解码吞吐量提升了3.56×。我们的代码可在该https URL获取。
英文摘要
Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches. Existing sparse attention reduce computation by attending to fewer KV pairs, but often suffer from substantial accuracy degradation, require additional training, or rely on expensive hashing. In this work, we present BinaryPC, a training-free, data-aware hashing-based sparse attention for long-context LLMs. BinaryPC constructs compact binary hash codes and corresponding hash function by computing binary principal components of data. Unlike Locality-Sensitive Hashing (LSH) with data-independent random projections or learned non-linear hashing methods, BinaryPC constructs binary codes that explicitly preserve the structural information of data without requiring gradient-based training. Comprehensive experiments across multiple model families and long-context benchmarks show that BinaryPC preserves accuracy relative to full attention while achieving superior performance among sparse and hashing-based baselines. On modern GPUs, BinaryPC improves end-to-end decoding throughput by 3.56$\times$ over the FlashAttention kernel. Our code is available at https://github.com/yudaohai666/BPC.
CommentsICML 2026