arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型(LLM)无训练低秩压缩中的校准与截断误差传播研究

Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs

Mohanad Odema, Gabrielle De Micheli, Dayin Gou, Nilesh Malpeddi, Prathamesh Vaste, Jacob Song

arXiv 2608.08506首次发表:更新:

发表机构

LG Electronics, North America(北美LG电子)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对现有LLM无训练低秩压缩框架的两个关键局限,提出带校准修正的逐层压缩与带秩分配修正的迭代压缩方法,在Llama、Qwen3模型的零样本任务上实现1-2.5个准确率点提升。

AI 中文摘要

无训练低秩压缩框架因能有效减少模型参数数量同时保持任务级准确率,在大型语言模型(LLM)压缩领域日益受到关注。然而,现有最优(SOTA)框架存在两个关键局限:其一,校准数据激活值中的残差误差会在压缩过程中跨层累积,导致压缩时模拟的表示与推理时实际使用的表示出现失配;其二,压缩后层重要性分布得以保留的假设不成立。这两种效应共同导致压缩过程与部署模型间出现失配。本研究针对这些效应提出一种与现有框架兼容的简单无训练方法,包含:(1)带校准修正的逐层压缩;(2)带秩分配修正的迭代压缩。该方法在现有最优分解框架基础上实现,于Llama和Qwen3模型上,在各类基准测试及不同压缩率下评估,与按权重分解及联合分解基线相比,在零样本任务上展现出最高约1至2.5个准确率点的提升。

英文摘要

Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy. However, existing SOTA frameworks share two key limitations: (1) residual errors in calibration data activations accumulate across layers during compression, causing misalignment between representations simulated at compression time and those experienced at inference; (2) the assumption that layer importance distribution is preserved post-compression does not hold. Together, these two effects introduce misalignment in the compression process in relation to the deployed model. We study these effects and propose a simple, training-free methodology compatible with existing frameworks to mitigate them, comprising: (1) Layer-by-Layer Compression with Calibration Correction; (2) Iterative Compression with Rank Allocation Correction. Implemented atop an existing SOTA decomposition framework, and evaluated on Llama and Qwen3 models across various benchmarks and compression rates, our approach demonstrates up to ~1-2.5 accuracy point improvements over per-weight and joint decomposition baselines on zero-shot tasks.

CommentsCOLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑