AI 中文总结
该研究提出联合稀疏性、量化、低秩近似的压缩三位一体框架,通过MKOR、SLoPe等技术实现LLM高效压缩,提升精度与效率,优于现有方法及未压缩密集模型。
AI 中文摘要
高昂的计算成本和环境成本阻碍了大语言模型(LLM)的规模化部署。传统压缩技术(稀疏性、量化、低秩近似)通常单独应用,每种技术都存在精度-效率瓶颈。本论文提出了“压缩三位一体”,这是一种联合应用三大支柱的统一框架:稀疏性用于减少计算量,量化用于最小化内存带宽,低秩近似用于恢复精度。为加速预训练,我们将三位一体应用于优化器和模型架构。MKOR通过块对角稀疏性和低秩逆近似曲率,为量化状态保持数值稳定性;它将曲率更新复杂度从$O(d^3)$降至$O(d^2)$,并比KFAC最多加速1.85倍收敛。SLoPe通过针对N:M稀疏性的双剪枝反向传播,在训练的最后1%阶段使用低秩“惰性”适配器恢复精度,最多可加速1.25倍训练。对于训练后压缩,OPTIMA通过将权重重构表述为全局最优逐列二次规划,在零训练机制下稳定静态掩码,最多提升3.97%的零样本精度。若存在微调预算,PATCH通过学习0%至50%之间的动态混合稀疏率突破静态掩码的上限,实现最多1.38倍加速。最后,SLiM一次性实现完整的压缩三位一体,使用数学推导的低秩适配器恢复量化和稀疏性造成的信息损失,比现有最优方法最多提升5.66%精度,且在相同参数预算下比未压缩的密集模型表现高出0.6%。这些结果共同表明,联合应用压缩三位一体对于高效、规模化、高性能的LLM至关重要。
英文摘要
Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the "Compression Trinity," a unified framework that applies the three pillars jointly: sparsity to reduce computation, quantization to minimize memory bandwidth, and low-rank approximations to recover accuracy. To accelerate pretraining, we apply the Trinity to the optimizer and model architecture. MKOR approximates curvature via block-diagonal sparsity and low-rank inversion, maintaining numerical stability for quantized states; it reduces curvature update complexity from $O(d^3)$ to $O(d^2)$ and accelerates convergence by up to 1.85x over KFAC. SLoPe accelerates training by up to 1.25x via a double-pruned backward pass for N:M sparsity, using low-rank "lazy" adapters in the final 1% of training to recover accuracy. For post-training compression, OPTIMA stabilizes static masks in a zero-training regime by formulating weight reconstruction as globally optimal column-wise quadratic programs, improving zero-shot accuracy by up to 3.97%. Given a fine-tuning budget, PATCH breaks the ceiling of static masks by learning a dynamic hybrid sparsity ratio between 0% and 50%, yielding up to 1.38x speedups. Finally, SLiM realizes the full Compression Trinity in one shot, using mathematically derived low-rank adapters to recover information lost to quantization and sparsity, improving accuracy by up to 5.66% over state-of-the-art methods and outperforming uncompressed dense models at equal parameter budgets by 0.6%. Together, these results show that jointly applying the Compression Trinity is essential for efficient, scalable, high-performance LLMs.
CommentsPhD thesis, University of Toronto, 2026. 156 pages. Chapters extend MKOR (arXiv:2306.01685), SLoPe (arXiv:2405.16325), OPTIMA (arXiv:2512.13886), PATCH (arXiv:2509.23410), and SLiM (arXiv:2410.09615). Official record: https://utoronto.scholaris.ca/items/2cde1f98-6084-46b9-aae2-dbcd045f1215