发表机构
Ningbo Institute of Digital Twin, Eastern Institute of Technology; Munich Center for Machine Learning, LMU Munich(宁波东方理工大学宁波数字孪生研究院; 慕尼黑大学慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WIDE是首个端到端可微令牌级动态宽度剪枝框架,通过两阶段训练与剪枝-内核协同设计,在50%稀疏度下实现显著推理加速与更优精度保持,性能优于现有动态剪枝方法。
AI 中文摘要
剪枝是提升大语言模型(LLM)效率的可行方法。现有静态结构化剪枝方法对硬件友好,可实现实用的吞吐量提升,但其与输入无关的计算分配在高稀疏度下常导致显著的精度下降。近期的动态稀疏度方法通过适配各输入的计算量提升了质量保持能力,但大多仍局限于粗粒度的结构决策,在实际推理场景中的实际加速效果仍具挑战性。为解决这些问题,我们提出WIDE,这是首个专为预填充(prefill)和解码(decode)场景设计的端到端可微令牌级动态宽度剪枝框架。WIDE允许每个令牌动态选择注意力头组和前馈网络(FFN)通道组,从而实现细粒度的计算分配,将动态剪枝从层级决策扩展至神经元块级粒度。通过两阶段训练流程,WIDE学习到有效的令牌级稀疏执行模式,且相比现有方法实现了显著更优的质量保持。为使这种细粒度动态剪枝具备实用性,我们进一步提出剪枝-内核协同设计框架,将动态稀疏度加速分解为掩码重排序、与硬件无关的块级跳过以及与硬件相关的块内跳过,从而实现不同粒度下的高效执行。在50%稀疏度下,WIDE在仅校准设置下相比最先进的动态深度剪枝方法提供了55.1%的性能提升;在预填充和解码推理负载下,WIDE实现了接近理论值的内核级加速,预填充达1.98倍、解码达4.95倍,端到端加速则分别达1.68倍和1.55倍。我们的代码可在此https URL获取。
英文摘要
Pruning is a promising approach for improving the efficiency of LLMs. Existing static structured pruning methods are hardware-friendly and can deliver practical throughput gains, but their input-agnostic computation allocation often causes substantial accuracy degradation under aggressive sparsity. Recent dynamic sparsity methods improve quality retention by adapting computation to individual inputs, yet they remain largely limited to coarse-grained structural decisions and their practical acceleration under real-world inference scenarios remains challenging. To address these challenges, we present WIDE, the first end-to-end differentiable token-level dynamic width pruning framework designed for both prefill and decode scenarios. WIDE enables fine-grained computation allocation by allowing each token to dynamically select attention-head groups and FFN-channel groups, extending dynamic pruning beyond layer-level decisions to neuron-block-level granularity. Through a two-stage training pipeline, WIDE learns effective token-wise sparse execution patterns and achieves substantially better quality retention than existing approaches. To make such fine-grained dynamic pruning practical, we further propose a pruning--kernel co-design framework that decomposes dynamic sparsity acceleration into mask reordering, hardware-agnostic block-level skipping, and hardware-dependent intra-block skipping, enabling efficient execution across different granularities. At 50% sparsity, WIDE provides 55.1% performance boost when compared to the state-of-the-art dynamic depth pruning under calibration-only settings. Under prefill and decoding inference workloads, WIDE achieves close-to-theoretical kernel-level speedups of up to 1.98x for prefill and 4.95x for decoding, as well as 1.68x and 1.55x end-to-end acceleration. Our code is available at https://github.com/EIT-NLP/LLM-Pruning/tree/main/WIDE.
Comments30 pages, 19 figures