零数据自对弈预训练
Self-Play Pretraining with Zero Data
浏览论文内容
中文总结 AI 辅助
提出零数据自对弈预训练方法,通过生成器与学习器协同训练,在通用图灵机空间中搜索数据,实现自然数据零样本性能随计算量可预测提升。
中文摘要 AI 辅助
语言建模的进展一直由在越来越多的数据上进行预训练的规模扩展所驱动。然而,训练数据在很大程度上仍然是为模型而精心策划的。一种更通用的预训练方法将让模型学会生成对其自身改进最有用的数据。这将提供一种实际上无界的训练数据来源,其限制在于计算而非人类知识。我们引入了零数据自对弈预训练,这是实现这一愿景的初步概念验证。我们的程序将合成数据生成视为对所有可计算结构空间的搜索,灵感来源于所罗门诺夫归纳。从随机初始化开始,两个模型协同学习:一个生成器提出由通用图灵机解释的程序,生成字节序列,而一个学习器自回归地预测这些字节序列。学习器使用标准交叉熵进行训练,而生成器则使用强化学习进行训练,以产生处于学习器能力前沿的序列,从而产生自适应课程。通用图灵机为我们提供了对所有可计算数据生成过程的搜索空间,几乎不施加特定领域的结构,而自对弈则在该空间中搜索有用的训练数据。我们测试了自然数据上的零样本性能是否随自对弈计算量可预测地提高;这是对迁移的干净测试,因为生成器和学习器都没有在自然数据上训练。在多个自然数据集上,零样本损失表现出可预测的计算规模扩展。模型还表现出上下文学习能力,并在训练过程中发现了可识别的数学序列。
英文摘要
Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction. Starting from random initialization, two models learn in tandem: a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while a learner autoregressively predicts these byte sequences. The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum. A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data. We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training.
发表机构
- Stanford Institute for Theoretical Physics(斯坦福理论物理研究所)
- Tel Aviv University(特拉维夫大学)
- Stanford University(斯坦福大学)
- LAPTh, USMB(拉普拉斯理论物理实验室,萨瓦勃朗峰大学)
机构由 AI 辅助整理,请以论文原文为准。