AI 中文总结
研究针对长序列大语言模型推理计算成本高的问题,提出脉冲星注意力机制,用轻量级组件取代静态锚点,减少计算量,在多个模型测试中表现出色,优于星型注意力和密集注意力,有显著绝对增益。
AI 中文摘要
由于自注意力的二次复杂性,在长序列上使用大语言模型进行推理的计算成本很高。像星型注意力这样的分布式分块方法通过在主机间分片上下文来降低成本,但依赖于在每个主机前添加第一个块的静态、无内容副本。我们提出了脉冲星注意力机制,用两个轻量级、有内容感知的组件取代静态锚点:一个稳定softmax的小注意力汇聚前缀,以及通过选择包含全局罕见令牌的块的最大逆文档频率启发式构建的紧凑跨块摘要。这在保持相同的键值缓存占用的同时,将每个GPU的第一阶段浮点运算次数比星型注意力最多减少3.3倍。在RULER和BABILong以及Llama - 3.1 - 8B上,脉冲星注意力机制在序列长度达128K令牌时优于星型注意力和密集注意力,比密集基线绝对增益高达4.7%。
英文摘要
Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attention. Distributed blockwise methods such as Star Attention reduce this cost by sharding context across hosts, but rely on prepending a static, content-blind copy of the first block to every host. We propose Pulsar Attention, which replaces the static anchor with two lightweight, content-aware components: a small attention-sink prefix that stabilizes softmax, and compact cross-block summaries built via a Max-IDF heuristic that selects chunks containing globally rare tokens. This reduces the Phase 1 per-GPU FLOPs by up to 3.3x over Star Attention while retaining an identical KV cache footprint. On RULER with Llama-3.1-8B-Instruct, Pulsar Attention outperforms Star Attention at sequence lengths up to 128K tokens and remains competitive with dense attention across most tasks, with task-dependent absolute gains of up to 4.7% over the dense baseline.
Comments16 Pages