AI 中文总结
该研究提出谱权重衰减方法,在LLaMA等模型上可降低有效秩、提升压缩与推理速度,在含标签噪声的训练中也能提升分类准确率。
AI 中文摘要
标准权重衰减将每个权重矩阵视为向量,忽略其谱结构。我们提出谱权重衰减,一种后步解耦的核范数更新,采用加性而非乘性谱收缩。我们将该更新与近似近端下降关联,表明其在秩缺失附近对更新顺序的敏感性可超过传统的ℓ₂权重衰减。在参数规模为1.24亿至5亿的LLaMA模型上,谱权重衰减可降低有效秩,并在匹配验证损失下提升SVD-LLM压缩效果。在5亿参数、4%的失真预算下,其压缩率达1.89倍,GPU推理速度提升1.18倍,而标准权重衰减对应的数值分别为1.14倍和1.01倍。在含60%标签噪声的固定时域训练中,其在MNIST数据集上的最终平均干净测试准确率较匹配的ℓ₂正则化提升最高17.8个百分点,在四个BERT-base任务上提升最高4.6个百分点。代码可在该https URL获取。
英文摘要
Standard weight decay treats each weight matrix as a vector and ignores its spectral structure. We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage. We connect the update to approximate proximal descent and show that its sensitivity to update order can exceed that of conventional $\ell_2$ weight decay near rank deficiency. Across LLaMA models with $124$M to $500$M parameters, spectral weight decay lowers effective rank and improves SVD-LLM compression at matched validation loss. At $500$M and a $4\%$ distortion budget, it reaches $1.89\times$ compression and $1.18\times$ GPU inference speedup, compared with $1.14\times$ and $1.01\times$ after standard weight decay. Under fixed-horizon training with $60\%$ label noise, it also improves final mean clean-test accuracy over matched $\ell_2$ regularization by up to $17.8$ points on MNIST and $4.6$ points across four BERT-base tasks. Code is available at https://github.com/brain-lab-research/SpectralWD.
Comments20 pages, 12 figures, 5 tables