AI 中文总结
针对大语言模型参数规模增长需求,提出LACE-SVD框架,先依候选压缩率估算损失增加量分配秩预算,再用局部更新和误差校正优化模型,降低累积误差传播,效果优于现有方法。
AI 中文摘要
大语言模型参数规模快速增长,对高效压缩技术需求强烈。低秩压缩被广泛采用,但现有基于SVD的方法有局限。本文提出LACE-SVD,先估算候选层压缩率引起的校准负对数似然增加,解决预算受限分配问题,再用闭式局部更新优化并校正残差流输出模块,减少累积误差传播。实验表明该方法在高压缩率下效果良好。
英文摘要
The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques. As a hardware-agnostic and highly compatible approach, low-rank compression has been widely adopted to reduce both memory footprint and computational cost. However, existing SVD-based methods are still largely driven by local reconstruction objectives, overlooking two critical limitations: rank budgets are often allocated without explicitly considering layer-wise loss sensitivity, and local approximation errors can propagate and accumulate through the residual stream, leading to amplified global deviations from the original model. To address these issues, we propose LACE-SVD, a Loss-Aware SVD framework with Cumulative Error correction for LLM compression. LACE-SVD first estimates the calibration negative-log-likelihood increase induced by candidate layer-wise compression ratios and solves a budget-constrained allocation problem to assign rank budgets. It then refines the compressed model with closed-form local updates and introduces a propagation-aware correction for residual-stream output modules, reducing layer-output discrepancy as a proxy for cumulative error propagation. Experimental results demonstrate that at a high compression ratio (0.6), the WikiText-2 PPL of our method on LLaMA-7B (32.57) is significantly better than that of Dobi-SVD (46.18).
Comments12 pages, 5 figures, 5 tables