比Flash更快:利用注意力稀疏性实现高效长上下文解码
Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding
浏览论文内容
中文总结 AI 辅助
该研究提出硬件-算法协同设计框架FFD,利用注意力稀疏性实现长上下文解码的高效性,获内核级11.6倍加速、端到端吞吐量2.37倍提升,在基准数据集上保持精度。
中文摘要 AI 辅助
长上下文大语言模型(LLMs)的发展受限于解码阶段的内存带宽瓶颈及注意力机制的二次复杂度。为克服基于元数据的指标内存开销与自适应选择策略计算效率之间的固有权衡,我们提出Faster Flash Decoding(FFD),一种旨在突破长上下文解码内存墙的新型硬件-算法协同设计框架。FFD将选择器与计算单元整合为完全融合的内核,通过低比特量化的内容感知扫描替代外部元数据索引。此外,我们引入top-delta策略,动态过滤以实现分布自适应稀疏性,无需全局同步。FFD提供无需训练、即插即用的解决方案,还可复用扫描结果进行计算,实现最高11.6倍的内核级加速,扩展至256K上下文长度,端到端吞吐量提升2.37倍。在RULER与LongBench上的经验验证表明,FFD在保持模型精度的同时实现了高比例稀疏性,代码可在该URL获取。
英文摘要
The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding. To overcome the inherent trade-offs between the memory overhead of metadata-based metrics and the computational inefficiency of adaptive selection strategies, we present Faster Flash Decoding (FFD), a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding. FFD integrates the selector and computer into a fully fused kernel, replacing external metadata indices with content-aware scanning via low-bit quantization. Furthermore, we introduce the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization. Offering a training-free and plug-and-play solution, FFD also enables the reuse of scanning results for computation, achieving up to 11.6x kernel-level speedup and scaling to 256K context length, with 2.37x end-to-end throughput improvement. Empirical validation on RULER and LongBench confirms that FFD maintains model accuracy while delivering high-ratio sparsity, with code available at https://github.com/qluoluo/faster-flash-decoding
发表机构
- Fudan University(复旦大学)
- Shanghai Innovation Institute(上海创新研究院)
- Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。