发表机构
New York University; Carnegie Mellon University(纽约大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用自生成训练数据的顺序编码突破模型压缩极限,教师模型从学生分布选样本,学生代码记录选择,代码长度与参数和熵无关且更短,还能揭示现象、给出泛化保证、隔离可学习信息。
AI 中文摘要
压缩对于智能至关重要。能够将训练数据表示为短代码的模型发现了有助于泛化的规律。大型神经网络可能学习的函数比其参数数量所显示的要简单得多,但构建实现这种简单性的代码具有挑战性。基于参数的方法如量化产生的代码长度与模型大小成比例,对参数存储的信息量不敏感。顺序编码通过压缩训练轨迹绕过了这个问题,但对确切的数据序列进行编码,而不管模型学习了多少,当数据具有高熵时会产生大代码。我们引入顺序编码,其中教师模型从学生自身的分布中选择训练样本。学生的代码只记录这些选择,只有在教师和学生意见不一致时才花费比特。由此产生的代码长度与参数数量和数据熵无关,并且通常比顺序编码的对应物短几个数量级,优势随着规模的增加而增长。这种压缩揭示了先前压缩器无法触及的现象。在保持损失固定的情况下,尽管参数更多,更大的模型和集成压缩到更小的尺寸。将顺序编码插入 PAC - 贝叶斯界,为数十亿参数的语言模型产生了最先进的泛化保证,甚至在零误差的情况下也优于基于激进的训练后量化构建的界。在计算最优状态下,随着模型相对于数据集大小变得越来越可压缩,该界随着规模而收紧。相同的代码预测,当模型训练多个 epoch 时会逐渐过拟合。它还将数据集中可学习的信息与其不可预测的随机内容隔离开来,揭示出低熵文本比高熵图像数据拥有更多可学习的结构。
英文摘要
Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization. Large neural networks may learn functions far simpler than their parameter counts suggest, but it is challenging to construct codes that realize this simplicity. Parameter-based methods such as quantization produce code lengths that scale with model size, insensitive to how much information the parameters store. Prequential coding bypasses this issue by compressing the training trajectory, but codes the exact data sequence regardless of how much the model learns, yielding large codes when the data has high entropy. We introduce requential coding, where a teacher model selects training samples drawn from the student's own distribution. The student's code records only these selections, which cost bits only where teacher and student disagree. The resulting code length is independent of parameter count and data entropy, and often orders of magnitude shorter than the prequential counterpart, with an advantage that grows with scale. This compression sheds light on phenomena inaccessible to prior compressors. Holding loss fixed, larger models and ensembles compress to much smaller sizes despite more parameters. Plugged into a PAC-Bayes bound, the requential code yields state-of-the-art generalization guarantees for billion-parameter LLMs, outperforming bounds built on aggressive post-training quantization even granted zero error. The bound tightens with scale in the compute-optimal regime, as models become increasingly compressible relative to dataset size. The same code predicts that models gradually overfit when trained for multiple epochs. It also isolates the learnable information in a dataset from its unpredictable, random content, revealing that lower-entropy text holds far more learnable structure than higher-entropy image data.
CommentsCode available at https://github.com/shikaiqiu/requential-coding