arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoWindow Attention:完全因果覆盖是一种集体属性

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

arXiv 2609.32704首次发表:更新:

AI 中文总结

CoWA 通过 KV 头分布式窗口实现集体因果覆盖,减少冗余计算,性能接近 FullAttn,延迟大幅降低。

AI 中文摘要

FullAttn 反复将完整的因果历史暴露给每个注意力头,即使使用 IO 高效密集内核,也会产生大量冗余计算和内存流量。我们引入了 CoWA,一种结构化注意力架构,将因果历史的访问分布到 KV 头上。所有头共享近对角线和前缀汇窗口,而互补的长距离窗口划分剩余的历史。它们的并集提供了完整的因果覆盖,尽管每个头稀疏地关注远处的标记。这种位置定义的注意力模式不需要学习路由器或索引器,在训练和推理期间一致使用,并与 KV 头张量并行性对齐。在 8K 的窗口匹配消融中,隔离了互补长距离分配的效果:CoWA 具有 100% 的集体覆盖率达到 89.73% 的准确率,而 FullAttn 为 89.97%,而重复的长距离窗口表现明显更差。在更广泛的受控关联回忆比较中,使用匹配的标记预算,随着上下文增长,CoWA 紧密跟踪 FullAttn,而其他稀疏模式丢失了很大一部分关联。在 128K 标记的注意力算子基准测试中,使用张量并行性,CoWA 在训练期间将前向和后向延迟分别降低了 7.4 倍和 8.6 倍,推理期间解码延迟降低了 3.0 倍,相对于 FullAttn。其每个秩的峰值算子内存在训练期间与 FullAttn 匹配,解码期间低 7.6 倍。在从 0.6B 到 14B 参数的缩放定律训练中,CoWA 在困惑度上紧密跟踪 FullAttn,同时减少了总训练 FLOPs。由此产生的 14B 模型和从单独继续训练得到的 32B 模型在知识、推理和长上下文检索得分上与 FullAttn 相当。这些结果表明,完整的因果覆盖可以是头集合的集体属性,而不是每个头的重复属性。

英文摘要

FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. All heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. Their union provides full causal coverage although each head attends sparsely to distant tokens. This position-defined attention pattern requires no learned router or indexer, is used consistently during training and inference, and aligns with KV-head tensor parallelism. A window-matched ablation at 8K isolates the effect of complementary long-range allocation: CoWA with 100% collective coverage reaches 89.73% accuracy, compared with 89.97% for FullAttn, while duplicated long-range windows perform substantially worse. Across a broader controlled associative-recall comparison with matched token budgets, CoWA closely tracks FullAttn as the context grows, whereas other sparse patterns lose a substantial fraction of the associations. In an attention-operator benchmark at 128K tokens with tensor parallelism, CoWA reduces forward and backward latency during training by 7.4x and 8.6x and decoding latency during inference by 3.0x over FullAttn. Its per-rank peak operator memory matches FullAttn during training and is 7.6x lower during decoding. Across scaling-law training from 0.6B to 14B parameters, CoWA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results show that full causal coverage can be a collective property of the head ensemble rather than a duplicated property of every head.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑