发表机构
Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReScraper是一个0.6B参数统一模型,取代手写启发式规则完成网络数据抓取与清洗,通过四种操作提升LLM预训练语料质量,在多个规模上相对提升DCLM Core分数3.8-4.7%。
AI 中文摘要
LLM预训练语料库通常由一系列手写启发式规则进行清洗。一个启发式抓取器从HTML中提取主要内容,随后数十个基于规则的过滤器对其进行清洗,因此语料库质量受限于规则的粗糙程度和准确性。在本工作中,我们提出ReScraper,一个仅有0.6B参数的统一语言模型,取代了整个堆栈。为训练ReScraper,我们精心从三个教师模型的输出中整理监督数据,使其学会首先从原始数据中提取主要内容,然后在四种操作中选择:按提取结果保留页面、编辑掉噪声行和片段、完全删除页面,或在页面写得差但有信息量时进行重写。基于相同的爬取数据池,在我们整理的数据上预训练400M、1.4B和2.8B模型,在每种规模下相对于最强基线(包括昂贵的多智能体整理)将DCLM Core分数相对提升3.8--4.7%。我们的分析表明,每种操作都扮演着独特且互补的角色,并且在一个模型中同时进行提取和清洗优于级联的独立模型。ReScraper还将其操作集中在需要处理的页面上,最大程度地提高质量差的页面的质量,同时保持语料库的多样性。这些结果证明了AI4AI在预训练数据整理中的可行性和有效性,即一个小型学习模型接管了流水线中原本由手写启发式规则处理的整个阶段。我们在以下网址开源我们的代码:此https URL
英文摘要
LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a unified language model of only 0.6B parameters that replaces this entire stack. To train ReScraper, we carefully curate supervised data from the outputs of three teacher models, so it learns to first extract the main content from raw data and then choose among four operations: keeping the page as extracted, editing out noisy lines and spans, deleting it entirely, or rewriting it when it is poorly written but informative. Based on the same crawled data pool, pretraining 400M, 1.4B, and 2.8B models on our curated data improves the DCLM Core score by a relative 3.8--4.7% over the strongest baseline at each scale, including the costly multi-agent curation. Our analyses show that each operation plays a distinct and complementary role, and that extracting and cleaning in one model outperforms a cascade of separate models. ReScraper also concentrates its operations on the pages that need them, raising the quality of poor pages the most while keeping the corpus diverse. These results demonstrate the feasibility and effectiveness of AI4AI for pretraining data curation, where a small learned model takes over an entire stage of the pipeline from hand-written heuristics. We open-source our code at https://github.com/cxcscmu/ReScraper