arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

显而易见却常被忽视:规范元素对极端LLM稀疏性的重要意义

Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity

Hyeondo Jang, Kwanhee Lee, Dongyeop Lee, Namhoon Lee

arXiv 2609.06557首次发表:更新:

发表机构

POSTECH(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过渐进式稀疏化框架结合二阶显著性与协调持续训练,在LLaMA-2和Qwen-3模型上实现高达99%稀疏度下保持性能,显著超越现有方法。

AI 中文摘要

大型语言模型(LLMs)通常被认为在激进稀疏化下表现脆弱,维持可靠性能通常需要坚持适度的稀疏水平。然而,近期研究表明,LLMs对高稀疏度的韧性远超先前认知,这将该问题重新定义为设计挑战而非根本限制。在本工作中,我们通过重新审视在此规模下相对未被充分探索的基础剪枝策略,挑战了非结构化训练后LLM剪枝的感知极限。通过采用二阶显著性并结合与稀疏度进展协调的持续训练所构成的渐进式稀疏化框架,我们展示了预训练LLMs能够在远超常见研究稀疏度区间下保持强大性能。在LLaMA-2和Qwen-3模型家族中,我们的方法在高达99%稀疏度下提升了困惑度和下游任务准确率,超越了当前最先进水平及代表性基线。具体而言,在LLaMA-2-7B上,我们的方法在95%和99%稀疏度下分别实现了WikiText-2困惑度13.48和19.67,同时在95%稀疏度下提供了3.23倍解码加速和6.21倍内存节省。综合来看,我们的结果表明,LLMs可以被推向极端稀疏度同时保持强大性能,为在该领域进一步改进稀疏模型奠定了基础。

英文摘要

Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires sticking to moderate sparsity levels. However, recent studies suggest that LLMs are more resilient to high sparsity than previously thought, reframing the problem as a design challenge rather than a fundamental limitation. In this work, we challenge the perceived limits of unstructured post-training LLM pruning by revisiting elementary pruning strategies that have remained relatively underexplored at this scale. Through a progressive sparsification framework with second-order saliency and continued training coordinated with sparsity progression, we show that pretrained LLMs can retain strong performance far beyond commonly studied sparsity regimes. Across LLaMA-2 and Qwen-3 model families, our approach improves perplexity and downstream accuracy up to 99\% sparsity, surpassing both the current state-of-the-art and representative baselines. Precisely, on LLaMA-2-7B, our approach achieves WikiText-2 perplexities of 13.48 and 19.67 at 95\% and 99\% sparsity, respectively, while delivering 3.23$\times$ decoding speedup and 6.21$\times$ memory savings at 95\% sparsity. Taken together, our results show that LLMs can be pushed into extreme sparsity while retaining strong performance, providing a foundation for further improving sparse models in this regime.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑