Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression
压缩三位一体:探索用于大语言模型(LLM)压缩的稀疏性、量化与低秩近似
专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);pretraining(abstract)
AI总结 该研究提出联合稀疏性、量化、低秩近似的压缩三位一体框架,通过MKOR、SLoPe等技术实现LLM高效压缩,提升精度与效率,优于现有方法及未压缩密集模型。
Comments PhD thesis, University of Toronto, 2026. 156 pages. Chapters extend MKOR ( arXiv:2306.01685 (https://arxiv.org/abs/2306.01685) ), SLoPe ( arXiv:2405.16325 (https://arxiv.org/abs/2405.16325) ), OPTIMA ( arXiv:2512.13886 (https://arxiv.org/abs/2512.13886) ), PATCH ( arXiv:2509.23410 (https://arxiv.org/abs/2509.23410) ), and SLiM ( arXiv:2410.09615 (https://arxiv.org/abs/2410.09615) ). Official record: this https URL (https://utoronto.scholaris.ca/items/2cde1f98-6084-46b9-aae2-dbcd045f1215)