发表机构
KAUST(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
QK-Wanda通过耦合查询与键的删除成本进行非结构化剪枝,无需梯度更新,在多个大模型上降低QK重建误差并提升下游性能,同时揭示局部重建预测模型质量的局限。
AI 中文摘要
Wanda(Sun等人,2024)通过在每个线性投影内独立地对权重进行评分来剪枝大型语言模型,尽管查询和键通过点积相互作用。我们提出了QK-Wanda,它通过未掩蔽的预RoPE重建目标下的个体删除成本来对查询和键权重进行评分。它用来自相反投影的信息(键权重对应的查询信息,以及查询权重对应的键信息)增强了Wanda的评分,使得两个投影能够共享一个剪枝预算。其闭式评分不需要梯度或权重更新;在使用我们主要实验中的校准设置时,完整剪枝在A100上比Wanda慢1.3%,在H200上慢3.1%。我们评估了仅对QK进行剪枝在来自TinyLlama、Llama 2、Llama 3和Qwen2.5的15个模型上的表现,参数规模从0.5B到72B不等。相对于Wanda,QK-Wanda在50%稀疏度下平均将QK重建误差降低了60%,在80%稀疏度下降低了45%。下游收益取决于模型。在Llama 2 70B上80%稀疏度时,WikiText-2和C4困惑度分别降低了20.3%和13.5%,而平均零样本准确率上升了5.94个百分点。Qwen2.5-72B也有所改善,但Llama-3.1-70B尽管重建误差更低,困惑度却显著更高。这些结果既展示了耦合剪枝标准的潜力,也展示了局部重建作为模型质量预测器的局限性。
英文摘要
Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Llama 2, Llama 3, and Qwen2.5, spanning 0.5B-72B parameters. Relative to Wanda, QK-Wanda reduces QK reconstruction error by an average of 60% at 50% sparsity and 45% at 80%. Downstream gains depend on the model. At 80% sparsity on Llama 2 70B, WikiText-2 and C4 perplexity decrease by 20.3% and 13.5%, while mean zero-shot accuracy rises by 5.94 percentage points. Qwen2.5-72B also improves, but Llama-3.1-70B has substantially higher perplexity despite lower reconstruction error. These results show both the promise of coupled pruning criteria and the limits of local reconstruction as a predictor of model quality.
Comments81 pages, including appendices