超越改写:书籍级组织提升训练中合成教科书数据的质量
Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
浏览论文内容
中文总结 AI 辅助
本研究提出可扩展的合成教科书数据流程,发现书籍级组织可提升语言模型训练中期的下游性能,在Llama3-8B上也验证了该设计维度的有效性。
中文摘要 AI 辅助
合成教科书数据已提升了语言模型的预训练效果,但现有研究大多将其益处归因于生成内容或局部改写风格的特性。本研究探讨了一个不同的因素:相关内容是否被组织成连贯的书籍级文档。我们贡献了一个可扩展的合成流程,以及证明这种组织方式重要性的受控证据。该流程从预训练语料库中检索源材料,将其聚类为主题单元,规划分层目录,并将基于源材料的章节组装成完整书籍(我们的Full设置),生成了涵盖15000多个学科的68.6万本教科书(共320亿个token)。在训练中期混合数据中用该语料库替换自然书籍,可使下游性能平均提升+1.09。随后的受控比较分解了相关设计因素:内容匹配的Split条件保持生成文本和token不变,但将每章视为独立文档,Full设置的+1.02平均增益凸显了文档打包的作用;长度匹配的RandomConcat控制条件将不同书籍的章节拼接在一起,性能仍低于Full设置,排除了仅文档长度的影响;检索池匹配的Rephrase条件在相同受众-风格方案下独立改写单个检索文档,不进行聚类、目录规划或书籍组装,Full设置的+1.17增益证明了结构化合成的价值。在Llama3-8B模型上,Full设置同样优于RandomConcat和自然书籍,支持书籍级组织作为合成预训练数据设计的有用维度。
英文摘要
Synthetic textbook data has improved language model pre-training, but prior work largely treats the benefit as a property of generated content or local rewriting style. We study a different factor: whether related content is organized into coherent book-level documents. We contribute both a scalable synthesis pipeline and controlled evidence that this organization matters. The pipeline retrieves source material from a pre-training corpus, clusters it into topical units, plans hierarchical tables of contents, and assembles source-grounded sections into complete books (our Full setting), yielding 686K textbooks (32B tokens) across 15,000+ disciplines. Replacing natural books in a mid-training mix with this corpus improves downstream performance by +1.09 on average. Controlled comparisons then disentangle the relevant design factors. A content-matched Split condition holds generated text and tokens fixed but treats each section as an independent document; Full's +1.02 mean gain isolates document packaging. A length-matched RandomConcat control that joins sections from different books remains below Full, ruling out document length alone. A retrieval-pool-matched Rephrase condition independently rewrites individual retrieved documents under the same audience-by-style scheme, without clustering, TOC planning, or book assembly; Full's +1.17 gain demonstrates the value of structured synthesis. On Llama3-8B, Full likewise outperforms both RandomConcat and Natural Books, supporting book-level organization as a useful axis for synthetic pre-training data design.
发表机构
- Peking University(北京大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。