发表机构
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PHBA用top-k块稀疏检索和前缀状态替代滑动窗口注意力,在统一层内结合精确远程证据与压缩历史,实现高效长上下文建模与检索。
AI 中文摘要
结合线性序列模型与softmax注意力的混合架构,在高效长上下文建模与精确token检索之间提供了有效平衡。现有设计如原生混合注意力(NHA)将压缩的长期状态与滑动窗口注意力相结合,但其精确注意力局限于固定的局部窗口。在本工作中,我们引入了前缀状态混合块注意力(PHBA),该模型用top-k块稀疏检索取代局部滑动窗口注意力,并将每个检索到的块与一个紧凑的前缀状态耦合,该前缀状态概括了其前面的上下文。前缀状态由块边界处的门控线性递归构建,并与相应的token块一起被检索,使模型能够在统一层内将精确的远程证据与压缩的历史上下文相结合。我们进一步开发了一种硬件感知的Triton实现,该实现流式传输路由的token块和前缀状态,而无需物化大型中间张量。实验表明,与强线性及混合基线相比,PHBA在长上下文和检索性能上均有提升,同时保持了高效的训练和推理。
英文摘要
Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks, allowing the model to combine precise long-range evidence with compressed historical context within a unified layer. We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors. Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.