arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01662cs.AIcs.CLcs.DCcs.LG

LongCat 稀疏注意力:通过流感知分层跨层索引驯服闪电(Lightning)

LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing

Wen Zan, Jiaqi Zhang, Jianchao Tan, Hong Liu, Cunguang Wang, Xiang Li, Duyue Ma, Guanyu Wu, Yifan Lu, Fengcun Li, Yerui Sun, Peng Pei, Yuchen Xie, Xunliang Cai

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出软硬件协同设计的 LongCat 稀疏注意力框架,通过三种索引策略解决 DeepSeek 稀疏注意力的系统瓶颈,在多模型上实现与全注意力相当的性能,支撑 LongCat-2.0 开发并开源相关模型。

中文摘要 AI 辅助

DeepSeek 稀疏注意力(DSA)通过其闪电索引器(Lightning Indexer)实现高效的长上下文建模,但实际部署仍受限于索引器昂贵的 O(L²) 评分开销及其输出导致的硬件效率低下、不连续的内存访问模式。为解决这些系统级瓶颈,我们提出 LongCat 稀疏注意力(LSA),这是一个软硬件协同设计的框架,包含三个互补且正交的策略:(1)流感知索引,选择性地将分散的 KV 条目转换为硬件对齐的连续布局,以实现合并的高带宽内存(HBM)访问;(2)跨层索引,通过跨层蒸馏支持,复用单个层在后续层中产生的结果,从而摊销索引开销;(3)分层索引,采用从粗到细的评分方案,逐步缩小每个查询的候选集,从而大幅减少索引计算。从 69B-A3B 到 560B-A27B 模型的大规模扩展实验表明,LSA 在通用和长上下文基准上始终达到与全注意力相当的性能。此外,LSA 支持原生训练,上下文长度可达 100 万 token,并为 LongCat-2.0(1.6T-A48B)的开发提供支撑。为促进进一步研究,我们还引入并开源了 LongCat-Flash-Lite-Sparse(69B-A3B),它将 LSA 集成到 LongCat-Flash-Lite 中,并纳入更新的长上下文训练语料库。

英文摘要

DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuous memory-access patterns induced by its outputs. To address these system-level bottlenecks, we introduce LongCat Sparse Attention (LSA), a hardware-algorithm co-designed framework comprising three complementary and orthogonal strategies: (1) Streaming-Aware Indexing, which selectively converts scattered KV entries into hardware-aligned contiguous layouts to enable coalesced HBM access; (2) Cross-Layer Indexing, which amortizes indexing overhead by reusing the results produced by a single layer across consecutive layers, supported by cross-layer distillation; and (3) Hierarchical Indexing, which adopts a coarse-to-fine scoring scheme to progressively narrow the candidate set for each query, thereby substantially reducing indexing computation. Extensive scaling experiments, ranging from 69B-A3B to 560B-A27B models, demonstrate that LSA consistently achieves performance on par with full attention across both general-purpose and long-context benchmarks. Moreover, LSA supports native training with context lengths of up to one million tokens and underpins the development of LongCat-2.0 (1.6T-A48B). To facilitate further research, we also introduce and open-source LongCat-Flash-Lite-Sparse (69B-A3B), which integrates LSA into LongCat-Flash-Lite and incorporates an updated long-context training corpus.

发表机构

  • Meituan(美团)
  • LongCat Team(LongCat团队)

机构由 AI 辅助整理,请以论文原文为准。

↑