超越单维压缩:大语言模型的复合稀疏前沿
Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models
浏览论文内容
中文总结 AI 辅助
研究大语言模型压缩,提出复合稀疏框架,结合静态参数剪枝和动态令牌级计算,通过低秩近似、通道剪枝及引入轻量级路由器实现。实验表明其在相同总稀疏度下优于单机制压缩,揭示跨维度干扰及有效分配方式,为改进压缩提供实用途径。
中文摘要 AI 辅助
大语言模型通常通过静态参数剪枝或动态令牌级计算进行压缩,但过度稀疏化会在超过基本稀疏边界时导致性能迅速下降。本文探讨结合这两种机制能否通过分担压缩负担来延迟这种性能下降。研究了一个极简复合稀疏框架,先应用低秩近似和通道剪枝获得静态压缩主干,再引入轻量级路由器进行逐令牌动态层跳过。实验表明,在相同总稀疏度下,复合稀疏始终优于单机制压缩,延迟了理解任务的衰减点并保持更强建模性能。进一步分析揭示了参数剪枝和令牌跳过之间的跨维度干扰,并表明在固定稀疏预算下接近平衡的分配最有效。这些结果表明复合压缩为改进大语言模型压缩提供了实用方法,同时揭示了最终限制进一步压缩的更广泛跨维度稀疏边界。
英文摘要
Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary. This work asks \emph{whether combining these two mechanisms can delay such degradation by distributing the compression burden}. We study a minimalist compound sparsity framework that first applies low-rank approximation and channel pruning to obtain a statically compressed backbone, and then introduces lightweight routers for per-token dynamic layer skipping. This design enables independent control of parameter sparsity and token-level computation sparsity. Experiments across language understanding and modeling benchmarks show that compound sparsity consistently outperforms single-mechanism compression under the same total sparsity, delaying the decay point on understanding tasks and preserving stronger modeling performance. Further analysis reveals cross-dimensional interference between parameter pruning and token skipping, and shows that near-balanced allocation is most effective under a fixed sparsity budget. These results demonstrate that compound compression provides a practical way to improve LLM compression, while revealing a broader cross-dimensional sparsity boundary that ultimately limits further compression. Code will be available at https://github.com/EIT-NLP/LLM-Pruning.
发表机构
- Ningbo Institute of Digital Twin, Eastern Institute of Technology(宁波数字孪生研究院,东方理工大学)
机构由 AI 辅助整理,请以论文原文为准。