AI 中文总结
研究针对预训练数据处理未适应各示例需求的问题,提出DataOrchestra框架,统一处理操作并为每个示例编排特定管道,经实验验证该框架在多基准测试中表现良好,在数学持续预训练中有效且能减少计算。
AI 中文摘要
预训练数据处理对大语言模型的下游性能至关重要。然而,许多现有方法在语料库或领域级别定义固定处理策略并统一应用于多个示例,未适应每个示例的需求。我们提出了DataOrchestra框架,它统一不同处理操作并为每个示例编排特定示例的管道。给定一块预训练数据,编排器决定丢弃、不处理或清理它。对于要清理的块,它选择一个或多个下游操作,从编程编辑到不同形式的基于大语言模型的重写。对于每个重写步骤,它进一步生成具体指令,由相应下游工具模型执行。我们在由DataOrchestra处理的网络数据上从头开始预训练从0.5B到7B的模型,并在11个基准测试中观察到相对于单个数据处理方法稳定的平均增益。DataOrchestra在数学持续预训练中也有效,优于更强的处理基线,同时通过跳过不必要的下游操作减少处理计算。
英文摘要
Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each example. We propose DataOrchestra, a framework that unifies different processing operations and orchestrates an example-specific pipeline for each example. Given a chunk of pretraining data, an orchestrator decides whether to drop, untouch, or clean it. For a chunk to be cleaned, it selects one or more downstream operations, ranging from programmatic editing to different forms of LLM-based rewriting. For each rewriting step, it further generates a concrete instruction, which is executed by the corresponding downstream tool model. We pretrain models from 0.5B to 7B from scratch on web data processed by DataOrchestra and observe stable average gains over individual data-processing methods across 11 benchmarks. DataOrchestra is also effective for math continued pretraining and outperforms stronger processing baselines, while reducing processing compute by skipping unnecessary downstream operations.
Comments36 pages