徘徊与丧失:低比特语言模型中的知识坍缩
Linger and Lose: Knowledge Collapse in Low-Bit Language Models
浏览论文内容
中文总结 AI 辅助
本研究揭示低比特语言模型存在知识坍缩现象,即三值权重模型在训练中先获取知识容量后大量丧失,并定位到输出头不可预测值导致的权重失控,通过预热-稳定-衰减调度可显著提升容量。
中文摘要 AI 辅助
使用三值权重训练语言模型通常以损失和下游准确率来评判,这些指标相对于全精度仅显示出适度的代价。我们证明这些指标可能掩盖更大的失败。相反,我们衡量知识容量,即每个参数存储的事实比特数,使用具有已知信息内容的合成传记。我们从头训练GPT-2风格的模型,参数规模从2.5M到50M,涵盖五种精度。在标准余弦调度下,三值模型仅保留同等训练的fp16模型容量的6%。这一差距随模型规模扩大而增大,而困惑度仅上升1.4至1.6倍。在整个训练过程中测量,这些模型先获取容量,然后失去大部分容量。我们将这种知识坍缩识别为学习率停滞不稳定的现象。在无衰减的情况下保持在1--2×10^{-4}附近,坍缩前的模型在几百次暴露内坍缩,而恢复到安全速率并不能恢复容量。然后我们将坍缩定位在输出头。它发生在一个模型永远无法预测的值上。那里的权重无限制增长,而模型做出的其他每个预测均不变。我们发现使该值可预测即可消除坍缩,无论哪个属性携带它。采用预热-稳定-衰减调度并带有10%冷却时间,在25M规模下三值容量增加2.7倍,在50M规模下增加4.3倍。在更低精度下收益更大。降低输出头的学习率可防止从头训练时的失败。我们测试的后训练量化方法在4比特以下无法恢复可测量的容量。我们的发现建议在训练过程中以保留容量来评判低精度训练,而非最终损失。
英文摘要
Training language models with ternary weights is commonly judged by loss and downstream accuracy, which record only a modest cost relative to full precision. We show that these metrics can conceal a much larger failure. We instead measure knowledge capacity, the factual bits stored per parameter, on synthetic biographies with known information content. We train GPT-2-style models from scratch with 2.5M to 50M parameters at five precisions. Under the standard cosine schedule, ternary models retain as little as 6% of an identically trained fp16 model's capacity. The deficit widens with model size while perplexity rises by only 1.4 to 1.6 times. Measured throughout training, these models first acquire capacity and then lose most of it. We identify this knowledge collapse as a learning-rate dwell instability. Held near $1$--$2\times10^{-4}$ with no decay, a pre-collapse model collapses within a few hundred exposures, and returning to a safe rate does not restore capacity. We then locate the collapse in the output head. It happens at a value the model can never predict. The weights there grow unchecked, while every other prediction the model makes is unchanged. We find that making the value predictable removes the collapse, regardless of which attribute carries it. A warmup-stable-decay schedule with a 10% cooldown increases ternary capacity by 2.7 times at 25M and 4.3 times at 50M. Gains are larger at lower precision. Cutting the output head's learning rate prevents failure when training from scratch. The post-training quantization methods we tested recover no measurable capacity below 4 bits. Our findings suggest judging low-precision training by retained capacity during training rather than final loss.