arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11163cs.LGcs.CL

LILA:通过潜在谱几何实现大型语言模型的无校准结构化剪枝

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

Sankar Behera, Dhruv Singh, Anshika Agnihotri, Raj Kumar Choudhary, Satyadev Ahlawat, Yamuna Prasad

首次发表
浏览论文内容

中文总结 AI 辅助

LILA利用谱几何的KS距离评估神经元重要性,实现无需校准数据或训练的结构化剪枝,在零样本和微调后均超越现有方法,并揭示高压缩下的架构瓶颈。

中文摘要 AI 辅助

大型语言模型(LLM)的结构化剪枝提供了硬件高效的压缩方式,然而现有方法在剪枝时需要校准数据、梯度计算或大型辅助策略网络。LILA(潜在信息层分析)通过全连接前馈网络(FFN)权重矩阵的完整与神经元消融后的经验奇异值分布之间的科尔莫戈罗夫-斯米尔诺夫(KS)距离来评估神经元重要性,提供了一种无需训练、校准数据或辅助网络的闭式谱规则。在无需任何微调的情况下,LILA在LLaMA-2-7B上以25%稀疏度进行零样本准确率测试时,比PruneNet(4500万参数的强化学习策略)高出1.57个百分点,并在所有稀疏度水平上优于使用WikiText-2校准的SliceGPT,最高提升6.0个百分点,同时保持原始架构不变。经过一个周期的LoRA恢复微调后,LILA达到了极具竞争力的性能,在LLaMA-2-7B和Phi-2上与重度校准的SliceGPT基线相差仅0.48个百分点以内,尽管使用了零校准数据。神经正切核分析证实,与随机剪枝相比,功能失真减少了22倍,为谱重要性标准提供了理论基础。最后,将LILA扩展为通过KS分数动态分配稀疏度预算,在中等压缩下实现了最先进的生成保留,同时在高压缩率下揭示了基本的单层架构瓶颈。

英文摘要

Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix, providing a closed-form spectral rule requiring no training, calibration data, or auxiliary network. Without any fine-tuning, LILA surpasses PruneNet (45M-parameter RL policy) by 1.57~pp in zero-shot accuracy on LLaMA-2-7B at 25\% sparsity, and outperforms WikiText-2-calibrated SliceGPT by up to 6.0~pp across all sparsity levels, while preserving the original architecture. After one epoch of LoRA recovery fine-tuning, LILA achieves highly competitive performance, matching the heavily calibrated SliceGPT baseline to within a 0.48~pp margin across LLaMA-2-7B and Phi-2, despite using zero calibration data. A Neural Tangent Kernel analysis confirms a 22$\times$ reduction in functional distortion versus random pruning, providing theoretical grounding for the spectral importance criterion. Finally, extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.

发表机构

  • IIT Jammu(印度理工学院贾姆穆分校)
  • EC Bikaner(比卡内尔工程学院)

机构由 AI 辅助整理,请以论文原文为准。

↑