发表机构
Viettel AI; Viettel Group; Hanoi University of Science and Technology(越南电信人工智能公司; 越南电信集团; 河内科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出轻量级半结构化剪枝框架RoI,通过可微子集采样学习稀疏掩码,减少参数与内存开销,在Qwen2.5 LLM家族上实现高效且具竞争力的性能,为LLM高效部署提供实用路径。
AI 中文摘要
半结构化N:M稀疏性已成为加速大型语言模型(LLM)的实用方向。然而,现有的可学习掩码方法会产生大量参数和内存开销,限制了它们向大型模型和激进稀疏性 regime 的可扩展性。在本研究中,我们从平衡效率与可扩展性的角度重新审视半结构化剪枝。我们提出了重要性储备池(Reservoir of Importance,RoI),这是一种轻量级半结构化剪枝框架,通过可微子集采样学习稀疏性掩码。与现有对所有可行N:M模式建模完整分类分布的方法不同,RoI为稀疏性掩码学习引入了紧凑对数参数化,并执行无放回采样以选择掩码,从而将可训练参数从组合复杂度降低至O(M)。因此,RoI所需的可训练参数减少了1.5-8.75倍,内存成本显著降低,同时与硬件友好的稀疏性模式完全一致。对Qwen2.5 LLM家族(0.5-7B参数)多个规模的广泛评估表明,RoI在强大的内存效率、稳定性和对更激进N:M稀疏性模式的可扩展性方面实现了有竞争力的性能,为高效LLM部署提供了实用路径。
英文摘要
Semi-structured $N$:$M$ sparsity has emerged as a practical direction for accelerating large language models (LLMs). However, existing learnable-mask approaches incur substantial parameter and memory overhead, limiting their scalability to large models and aggressive sparsity regimes. In this work, we revisit semi-structured pruning from a perspective that reconciles efficiency with scalability. We propose Reservoir of Importance (RoI), a lightweight semi-structured pruning framework that learns sparsity masks through differentiable subset sampling. Unlike prior methods that model full categorical distributions over all feasible $N$:$M$ patterns, RoI introduces a compact-logit parameterization for sparsity mask learning and performs sampling without replacement to select masks, thereby reducing trainable parameters from combinatorial complexity to $\mathcal{O}({M})$. As a result, RoI requires 1.5-8.75$\times$ fewer learnable parameters and significantly lower memory cost, while remaining fully aligned with hardware-friendly sparsity patterns. Extensive evaluations across multiple scales of the Qwen2.5 LLM family (0.5-7B parameters) demonstrate that RoI achieves competitive performance with strong memory efficiency, stability, and scalability to more aggressive $N$:$M$ sparsity patterns, offering a practical path toward efficient LLM deployment.
CommentsAccepted as an EMNLP 2026 Main Conference paper