arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29499cs.LG

任务感知频谱剪枝:一种用于高效大语言模型推理的掩码混合框架

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma

首次发表
浏览论文内容

中文总结 AI 辅助

提出任务感知频谱剪枝(TASP)框架,通过校准模块级频谱描述符并构建任务特定掩码,在43%FLOP减少下保持Llama-3-70B性能97.7%,解码加速1.44倍。

中文摘要 AI 辅助

静态剪枝对每个提示施加一种稀疏结构,尽管推理、检索、生成、编码和翻译可能依赖于语言模型的不同部分。我们提出了任务感知频谱剪枝(TASP),这是一个训练后框架,它根据测量的任务特定消融效应校准模块级频谱描述符,在稀疏掩码构建过程中关闭分组查询注意力(Grouped-Query Attention)和SwiGLU依赖,并将每个用户轮次路由到一个编译好的掩码,该掩码在预填充和解码过程中保持固定。一个模块不相交的试点首先确定频谱信号在完全校准之前是否具有信息性。在所述回顾性操作规则下,该试点在评估的Llama-3-8B和Llama-3-70B检查点上通过,但拒绝Qwen2.5-1.5B,表明适用性依赖于模型而非普遍适用。在43%的活跃FLOP减少下,Llama-3-70B基准测试工具保留了密集BF16分数的97.7±0.2%。在与部署匹配的INT8权重/BF16计算运行时,在单个A100 80GB上,编译的稀疏路径相对于密集BF16保留了97.3±0.2%,并将解码延迟从45.2±0.4降低到31.3±0.4毫秒/令牌,产生1.44倍加速。因子化消融、不相交模块测试、编译的结构化基线、路由损坏研究以及明确的136 GPU小时校准审计进一步界定了这些增益的来源和操作范围。

英文摘要

Static pruning imposes one sparse structure on every prompt, even though reasoning, retrieval, generation, coding, and translation can depend on different parts of a language model. We introduce Task-Aware Spectral Pruning (TASP), a post-training framework that calibrates module-level spectral descriptors against measured task-specific ablation effects, closes grouped-query-attention and SwiGLU dependencies during sparse-mask construction, and routes each user turn to one compiled mask that remains fixed throughout prefill and decoding. A module-disjoint pilot first determines whether the spectral signal is informative before full calibration. Under the stated retrospective operating rule, the pilot passes on the evaluated Llama-3-8B and Llama-3-70B checkpoints but rejects Qwen2.5-1.5B, demonstrating that applicability is model-dependent rather than universal. At a 43% active-FLOP reduction, the Llama-3-70B benchmark harness retains 97.7 +/- 0.2% of the dense BF16 score. In the deployment-matched INT8-weight/BF16-compute runtime on a single A100 80GB, the compiled sparse path retains 97.3 +/- 0.2% relative to dense BF16 and reduces decode latency from 45.2 +/- 0.4 to 31.3 +/- 0.4 ms/token, yielding a 1.44x speedup. Factorized ablations, disjoint-module tests, compiled structured baselines, routing-corruption studies, and an explicit 136-GPU-hour calibration audit further delimit the source and operating regime of these gains

发表机构

  • Iowa State University(爱荷华州立大学)
  • Kalinga Institute of Industrial Technology(卡林加工业技术学院)

机构由 AI 辅助整理,请以论文原文为准。

↑