一切都是训练:一种全合成单阶段大语言模型配方
It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
- PleIAs
- Sorbonne Center for Artificial Intelligence(索邦人工智能中心)
- Sciences Po Médialab(巴黎政治学院媒体实验室)
- EPFL(瑞士洛桑联邦理工学院)
- CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt(维尔茨堡-施韦因富特应用技术大学CAIRO)
- Paris Dauphine-PSL(巴黎多芬纳-PSL大学)
- Lattice, ENS-PSL(巴黎高等师范学院-巴黎文理研究大学Lattice实验室)
- TU Munich, Munich Center for Machine Learning(慕尼黑工业大学慕尼黑机器学习中心)
- KU Leuven(鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出SYNTH,一个从维基百科文章生成的全合成单阶段训练语料库,将预训练、中期和后期训练合并,在同等计算量下优于过滤网络数据,实现高数据效率的通用模型训练。
AI中文摘要:
当前的预训练数据集来源于网络爬取,带有其所有问题,并且并非为支持中期和后期训练流程而设计——例如,它们包含很少的显式推理。因此,许多前沿实验室已开始开发自己的内部数据集,从最先进的模型出发,以增强其预训练数据混合,例如,用推理轨迹来解决冷启动问题。尽管这些数据集被证明有效,但没有一个是公开的,并且这种所谓的合成数据对语言模型(包括小型模型)的知识和技能获取的影响仍然知之甚少。我们提出SYNTH,这是第一个开源合成语料库,源自58,698篇维基百科文章,通过结构化放大精选的百科种子,将预训练、中期和后期训练合并为一个单一训练阶段。我们通过训练一系列模型来评估SYNTH:一个56M的微型模型(Monad),0.3B-0.6B的稠密模型(Baguettotron),以及一个13B/1B激活的MoE。在同等计算量下,SYNTH优于过滤后的网络数据,并且我们的模型与类似规模的开源权重基线保持竞争力。由于SYNTH是从有根据的段落反向翻译而来,SYNTH训练的模型实现了高事实精度,尽管训练token减少了10-140倍,记忆由种子语料库定向。这些结果表明,合成数据集,包括我们的SYNTH数据集,能够从一小部分训练数据中产生有竞争力的通用模型,从而在前沿推进时实现快速迭代。这些发现为显著提高数据效率的通用模型,以及在没有指令或对话数据可用的领域特定模型开辟了可能性。最后,我们在宽松许可下公开发布我们的SYNTH数据集和Baguettotron模型套件,从而支持开源语言模型开发。
英文摘要:
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present SYNTH, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate SYNTH by training a suite of models: a 56M tiny model (Monad), 0.3B-0.6B dense models (Baguettotron), and a 13B / 1B-active MoE. At iso-compute, SYNTH outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because SYNTH is back-translated from grounded passages, SYNTH-trained models achieve high factual precision despite 10-140x fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our SYNTH dataset, are capable of producing competitive generalist models from a fraction of the training data, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our SYNTH dataset and the suite of Baguettotron models under a permissive license, thus supporting open-source language model development.