AI 中文总结
LegoLM针对大语言模型全局权重共享的两种失效模式,提出三种无数据适配策略,在GPT-2 small和Mistral-7B上实现优于PTQ-8bit的压缩效果与精度,证实选择性替换是核心机制。
AI 中文摘要
我们提出了LegoLM,这是一种面向大语言模型的结构化权重共享压缩框架,其基础是对全局权重共享为何失效以及如何修复它的系统研究。我们确定了两种不同的失效模式:1. 分布不匹配:对于维度d≤2的向量块,具有异构权重尺度的Transformer层会产生与d线性增长的尺度不匹配惩罚,且无法通过增加K来解决,从而导致困惑度(perplexity)在这一范围内占主导;2. 异常值主导:对于标量块,约1/K的权重位于Lloyd-Max最外层决策阈值之外,无法被任何质心表示;它们的错误表示会在各层累积,导致灾难性的质量损失。LegoLM通过三种无数据适配来解决这两种失效模式:1. 标量块编码以消除与d线性相关的不匹配分量;2. 百分位选择性替换,即识别并逐字保留异常值权重;3. 对前几个和最后几个Transformer块的边界层保护。在GPT-2 small(1.24亿参数)和Mistral-7B上,LegoLM在Mistral-7B上实现4.41倍压缩时,困惑度仅下降+0.03%,在2.67倍压缩时仅下降-0.02%,在质量和压缩率上均优于PTQ-8bit。在LAMBADA和HellaSwag上的下游评估证实,K=64、p=99%的LegoLM在5.12倍压缩时,精度保持在噪声范围内,超过了PTQ-8bit的压缩率且精度与其匹配。我们进一步发现,异常值主导随模型规模增大而加剧:K=128时的全替换仅使GPT-2 small的困惑度下降+23%,却使Mistral-7B灾难性下降+1,134,279%,而p=99%的选择性替换可将两个模型的下降幅度均控制在15%以内。受控消融实验证实,选择性替换是核心机制:将其添加到每层K均值聚类中也能实现近乎无损的质量,与LegoLM的差距在0.02%以内。
英文摘要
We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why global weight sharing fails and how to fix it. We identify two distinct failure modes. Distributional mismatch: for vector blocks of dimension d <= 2, transformer layers with heterogeneous weight scales impose a scale-mismatch penalty that grows linearly with d and cannot be resolved by increasing K, producing perplexity in the millions.Outlier dominance: for scalar blocks, a fraction ~1/K of weights lies beyond the outermost Lloyd-Max decision threshold and cannot be represented by any centroid; their misrepresentation accumulates across layers, causing catastrophic quality loss. \LegoLM{} resolves both failure modes via three data-free adaptations: 1 scalar-block encoding to eliminate the $d$-linear mismatch component, 2 percentile-selective replacement that identifies and preserves outlier weights verbatim, and 3 boundary-layer protection for the first and last transformer blocks. Across GPT-2 small (124M) and Mistral-7B, \LegoLM{} achieves +0.03% PPL degradation at 4.41X compression on Mistral-7B - outperforming PTQ-8bit in both quality and compression ratio - and -0.02% at 2.67X. Downstream evaluation on LAMBADA and HellaSwag confirms that \LegoLM{} at K=64, p=99% preserves accuracy within noise at 5.12 X compression, exceeding PTQ-8bit's compression ratio while matching its accuracy. We further discover that outlier dominance grows with model scale: full replacement at K=128 degrades GPT-2 small by only +23% but catastrophically degrades Mistral-7B by +1,134,279%, while selective replacement at p=99% rescues both models to under +15%. A controlled ablation confirms that selective replacement is the dominant mechanism: adding it to per-layer K-means also yields near-lossless quality, matching \LegoLM{} within 0.02%.
Comments8 pages, 3 figures