发表机构
University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
IrekoGPT通过保留投影矩阵并逐层校准,将预训练大语言模型转化为推理时可调宽度的可瘦身模型,在Llama和Qwen上优于PCA方法,尤其在高压缩率下。
AI 中文摘要
我们提出了IrekoGPT,一种事后方法,用于将预训练的大语言模型转换为可在推理时调整宽度的可瘦身模型。基于SliceGPT,我们保留其投影矩阵而不进行剪枝,使得单个模型能够以不同宽度暴露嵌套子网络。我们通过在多种压缩比率下校准每一层来提高鲁棒性,并通过无梯度的岭回归修正下游线性层。在Llama和Qwen模型上,初步结果显示相较于基于PCA的朴素瘦身方法有所改进,且在高压缩率下增益最大。代码可在以下网址获取:此https URL
英文摘要
We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at https://github.com/aimagelab/IrekoGPT
CommentsAccepted at the NeurIPS 2026 Workshop "AXIOM: Foundations of Efficient Deep Learning". 8 pages, 5 figures