arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06630cs.LG

稀疏低语者

The Sparsity Whisperer

Linghao Kong, Inimai Subramanian, Micah Adler, Dan Alistarh, Dan Gutfreund, Nir Shavit

AI总结:

本文针对大语言模型剪枝忽略输出差异的问题,提出Wisp、Wisp+、Whisper等差异感知剪枝方法,在Llama系列模型上提升了剪枝效果,可与其他技术结合优化性能。

AI中文摘要:

剪枝可降低大语言模型的推理成本,但现有剪枝准则主要保留大激活值或重构层输出。本文指出,现有研究忽略了MLP的上投影门和门控投影中对稀疏性敏感的神经元所执行的关键计算:将相似输入分离为不同输出。这表明有效剪枝不仅应保留激活值,还应更广泛地保留输出之间的差异。本文引入一系列基于该原理的差异感知剪枝方法:Wisp是一种一阶、无需更新的方法,通过输入差范数对权重评分;Wisp+会利用每个神经元分离最强的输入对,按神经元细化该评分;Whisper则是二阶方法,采用轻度正则化的差异海森矩阵作为重构目标。在参数规模7B至405B的Llama 2和3.1模型上,本文的二阶变体始终优于强大的基于重构的基线,无需更新的变体则优于感知激活的基线,尤其在受限设置中表现突出。对Wanda和SparseGPT的改进还延伸至结构化稀疏性、下游评估及其他模型家族。将本文的差异感知准则与RIA、ALPS等更强技术结合,可进一步提升性能,以可忽略的额外成本拓宽整体准确率-运行时间前沿。这些结果表明,保留输出差异是后训练LLM稀疏化中广泛有用且可组合的信号。

英文摘要:

Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. This suggests that effective pruning should preserve not only activations, but also the differences between outputs more broadly. We introduce a family of difference-informed pruning methods built upon this principle. Wisp is a first-order, update-free method that scores weights using input-difference norms, and Wisp+ refines this score neuronwise using the input pairs each neuron separates most strongly. Finally, Whisper is a second-order method that uses a lightly regularized difference Hessian as its reconstruction objective. Across Llama 2 and 3.1 models from 7B to 405B parameters, our second-order variant consistently improves over strong reconstruction-based baselines, while our update-free variants improve over activation-aware baselines, especially in constrained settings. The improvements over Wanda and SparseGPT extend to structured sparsity, downstream evaluations, and other model families. Augmenting stronger techniques such as RIA and ALPS with our difference-informed criteria yields further improvements, shifting the overall accuracy-runtime frontier outward at negligible additional cost. These results suggest that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.

补充信息

↑