发表机构
University of Oregon; Adobe Research; Hanoi University of Science and Technology(俄勒冈大学; 奥多比研究院; 河内国家大学自然科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长上下文 LLM 注意力预填充的二次缩放瓶颈,提出 CRISP 方法,通过结构代理路由和感知汇阈值消除噪声,在多基准上实现最优稀疏注意力性能及显著加速。
AI 中文摘要
长上下文大语言模型(LLM)推理的注意力预填充阶段呈二次方缩放,使得自注意力成为严重的计算瓶颈。传统稀疏注意力方法通过固定模式或离线分析缓解该问题,但缺乏适应输入依赖的注意力结构的灵活性。近期的动态方法通过实时将注意力头路由至稀疏模式解决该问题,但依赖间接的路由代理,存在开销,且预算分配机制忽略了 softmax 后的质量层次结构。我们提出 CRISP(Cliff-awaRe Input-adaptive Sparse Prefilling,考虑 Cliff 的输入自适应稀疏预填充),识别并解决该动态路由范式中的两个结构挑战。第一,我们证明路由决策可直接从代理注意力图的结构中读取。我们用 C_struct 替代 Jensen-Shannon 散度(JSD)路由,C_struct 是一种结构代理,测量垂直斜线兼容位置的质量,可复现 JSD 的路由决策,同时消除了池化矩阵乘法及后续 KL 散度的开销。第二,我们形式化 softmax 后的质量 Cliff,并从理论上证明严格累积覆盖阈值在长上下文下会累积 O(n) 背景噪声。CRISP 通过基于噪声基底的感知汇阈值解决该问题。实验中,在两个模型家族的 InfiniteBench、RULER 和 LongBench 上,CRISP 是整体表现最强的稀疏方法,在检索密集型基准上与精确的密集注意力相当或更优,在检索任务上较基线提升达 28.0 个百分点,且在 512k token 时实现最高达 5.30 倍的注意力加速,这主要得益于我们在选择过程中消除 O(n) 噪声的同时保留了结构完整性。
英文摘要
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe computational bottleneck. Traditional sparse attention methods mitigate this through fixed patterns or offline profiling, but lack the flexibility to adapt to input-dependent attention structure. Recent dynamic methods address this by routing heads to sparse patterns in real-time, but rely on indirect routing proxies with overhead and budget allocation mechanisms that overlook the post-softmax mass hierarchy. We present CRISP (Cliff-awaRe Input-adaptive Sparse Prefilling), which identifies and addresses two structural challenges in this dynamic routing paradigm. First, we show that the routing decision can be read directly off the structure of the proxy attention map. We replace the Jensen-Shannon Divergence (JSD) routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead. Second, we formalize the post-softmax mass cliff and demonstrate theoretically that strictly cumulative coverage thresholds accumulate O(n) background noise at long contexts. CRISP navigates this via a sink-aware threshold grounded in the noise floor. Empirically, across InfiniteBench, RULER and LongBench on two model families, CRISP is the strongest sparse method overall and matches or exceeds exact dense attention on retrieval-heavy benchmarks, recovering up to +28.0 pp on retrieval tasks over baselines and achieving up to a 5.30x attention speedup at 512k tokens, driven primarily by our O(n) noise elimination during selection while preserving structural integrity.
CommentsAccepted to EMNLP 2026 (Main Conference). 16 pages, 7 figures, 12 tables