arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

逐矩阵最优性不足:低秩大语言模型压缩的三级优化

Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

Huicheng Zhang, Xiyao Feng, Ze-Tong Li, Chengkai Zhu, Xiao Shi, Xiwei Pan, Jinguo Liu, Ge Bai, Xin Wang

arXiv 2609.15838首次发表:更新:

发表机构

Hong Kong University of Science and Technology (Guangzhou); QudeLeap Research(香港科技大学(广州); QudeLeap研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低秩LLM压缩,提出从矩阵到块再到模型的三级优化链,在LLaMA-7B上显著降低困惑度,块级阶段起正则化作用,但下游准确率仍不及稠密模型。

AI 中文摘要

逐矩阵奇异值分解(SVD)截断在白化Frobenius范数下是Eckart-Young最优的,但独立压缩矩阵的误差会通过块的非线性前向传播而累积。部分受量子多体方法中分层变分优化的启发,我们引入了一个三级链,将优化范围从单个矩阵扩展到Transformer块,再到整个模型:白化SVD(L1)、块级联合优化(L2)和端到端语言建模损失精炼(L3),全部仅使用256条校准序列,无需指令或恢复数据。在LLaMA-7B上以60%压缩率,该链将WikiText-2困惑度从42.1降至19.1,再降至11.4。块级阶段充当正则化器:跳过它会使Penn Treebank(PTB)困惑度恶化24个点,这一差距在我们的实验中额外的端到端训练未能弥合。困惑度的提升在20-80%压缩率、五种架构(最大13B参数)以及分布内和分布外基准上均保持,尽管跨架构行使用架构特定配置,且比率扫描未在统一协议下运行。使用更多校准数据时,跳过块级阶段变得具有竞争力,揭示了离线计算与数据之间的权衡。因此,我们仅声称在困惑度和压缩保真度上的改进;下游准确率仍远低于稠密模型。

英文摘要

Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius norm, but errors from independently compressed matrices compound through the block's nonlinear forward pass. Inspired in part by hierarchical variational optimization in quantum many-body methods, we introduce a three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model: whitened SVD~(L1), block-level joint optimization~(L2), and end-to-end language-modeling loss refinement~(L3), all from 256 calibration sequences, with no instruction or recovery data. On LLaMA-7B at 60% compression, the chain reduces WikiText-2 perplexity from 42.1 to 19.1 to 11.4. The block-level stage acts as a regularizer: skipping it worsens Penn Treebank (PTB) perplexity by 24 points, a gap that additional end-to-end training did not close in our experiments. Perplexity gains hold across 20-80% compression, five architectures up to 13B parameters, and both in-distribution and out-of-distribution benchmarks, though the cross-architecture rows use architecture-specific configurations and the ratio sweep was not run under one common protocol. With more calibration data, skipping the block-level stage becomes competitive, revealing an offline compute--data trade-off. We therefore claim improvements only in perplexity and compression fidelity; downstream accuracy remains well below the dense model.

Comments16 pages, to appear at EMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑