LongCat 稀疏注意力:通过流感知分层跨层索引驯服闪电(Lightning)
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
浏览论文内容
中文总结 AI 辅助
该研究提出软硬件协同设计的 LongCat 稀疏注意力框架,通过三种索引策略解决 DeepSeek 稀疏注意力的系统瓶颈,在多模型上实现与全注意力相当的性能,支撑 LongCat-2.0 开发并开源相关模型。
中文摘要 AI 辅助
DeepSeek 稀疏注意力(DSA)通过其闪电索引器(Lightning Indexer)实现高效的长上下文建模,但实际部署仍受限于索引器昂贵的 O(L²) 评分开销及其输出导致的硬件效率低下、不连续的内存访问模式。为解决这些系统级瓶颈,我们提出 LongCat 稀疏注意力(LSA),这是一个软硬件协同设计的框架,包含三个互补且正交的策略:(1)流感知索引,选择性地将分散的 KV 条目转换为硬件对齐的连续布局,以实现合并的高带宽内存(HBM)访问;(2)跨层索引,通过跨层蒸馏支持,复用单个层在后续层中产生的结果,从而摊销索引开销;(3)分层索引,采用从粗到细的评分方案,逐步缩小每个查询的候选集,从而大幅减少索引计算。从 69B-A3B 到 560B-A27B 模型的大规模扩展实验表明,LSA 在通用和长上下文基准上始终达到与全注意力相当的性能。此外,LSA 支持原生训练,上下文长度可达 100 万 token,并为 LongCat-2.0(1.6T-A48B)的开发提供支撑。为促进进一步研究,我们还引入并开源了 LongCat-Flash-Lite-Sparse(69B-A3B),它将 LSA 集成到 LongCat-Flash-Lite 中,并纳入更新的长上下文训练语料库。
英文摘要
DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuous memory-access patterns induced by its outputs. To address these system-level bottlenecks, we introduce LongCat Sparse Attention (LSA), a hardware-algorithm co-designed framework comprising three complementary and orthogonal strategies: (1) Streaming-Aware Indexing, which selectively converts scattered KV entries into hardware-aligned contiguous layouts to enable coalesced HBM access; (2) Cross-Layer Indexing, which amortizes indexing overhead by reusing the results produced by a single layer across consecutive layers, supported by cross-layer distillation; and (3) Hierarchical Indexing, which adopts a coarse-to-fine scoring scheme to progressively narrow the candidate set for each query, thereby substantially reducing indexing computation. Extensive scaling experiments, ranging from 69B-A3B to 560B-A27B models, demonstrate that LSA consistently achieves performance on par with full attention across both general-purpose and long-context benchmarks. Moreover, LSA supports native training with context lengths of up to one million tokens and underpins the development of LongCat-2.0 (1.6T-A48B). To facilitate further research, we also introduce and open-source LongCat-Flash-Lite-Sparse (69B-A3B), which integrates LSA into LongCat-Flash-Lite and incorporates an updated long-context training corpus.
发表机构
- Meituan(美团)
- LongCat Team(LongCat团队)
机构由 AI 辅助整理,请以论文原文为准。