发表机构
College of Computer Science, Chongqing University; School of Information Science and Engineering, Chongqing Jiaotong University(重庆大学计算机学院; 重庆交通大学信息科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RouteSparse为预算约束长上下文预填充提出输入条件稀疏模式路由方法,在Llama 3.1-8B-Instruct上实现6.5倍密集预填充速度,仅RULER指标降0.2点,平衡了质量与延迟。
AI 中文摘要
动态稀疏注意力可在不改变模型权重的情况下降低长上下文预填充的二次方成本。MInference为每个注意力头离线分配一种模式,并为每个提示估计该模式的稀疏索引。该设计高效,但假设注意力头的偏好模式和稀疏预算在不同输入间保持适用。我们提出RouteSparse,它将每个注意力头和提示段在小型GPU高效稀疏模式库中进行路由。一个低成本探测模块估计模式效用和不确定性;延迟感知路由器随后选择模式和预算,而不确定的情况则回退到更密集的掩码。我们将路由表述为约束风险最小化问题,从省略的概率质量中推导注意力输出误差的保证,并在长上下文检索、问答、摘要和语言建模任务上评估该方法。在Llama 3.1-8B-Instruct、128K token提示的设置下,RouteSparse实现了6.5倍于密集预填充的速度,同时相对于密集注意力的RULER指标下降0.2个点;而固定每个注意力头路由的方案则实现7.3倍速度,但下降1.6个点。 ablation实验证实,输入条件路由、硬件 profiling和选择性密集回退均对质量-延迟权衡有贡献。
英文摘要
Dynamic sparse attention can reduce the quadratic cost of long-context prefilling without changing model weights. MInference assigns each attention head one pattern offline and estimates that pattern's sparse indices for every prompt. This design is efficient, but it assumes that a head's preferred pattern and sparsity budget remain suitable across inputs. We introduce RouteSparse, which routes each head and prompt segment among a small library of GPU-efficient sparse patterns. A low-cost probe estimates pattern utility and uncertainty; a latency-aware router then selects a pattern and budget, while uncertain cases fall back to a denser mask. We formulate routing as constrained risk minimization, derive an attention-output error certificate from omitted probability mass, and evaluate the method on long-context retrieval, question answering, summarization, and language modeling. On Llama 3.1-8B-Instruct with 128K-token prompts, RouteSparse achieves $6.5\times$ dense prefill speed with a 0.2-point RULER drop relative to dense attention, compared with $7.3\times$ speed and a 1.6-point drop for fixed per-head routing. Ablations confirm that input-conditional routing, hardware profiling, and selective dense fallback each contribute to the quality--latency tradeoff.