发表机构
Peking University; Nanjing University(北京大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出REER-PT框架,通过逆向工程推理生成预训练数据的推理注释以增强数据,使模型在多个知识和推理基准上性能提升最高达2.07个百分点。
AI 中文摘要
随着语言模型的计算规模不断扩大,高质量训练数据正成为日益重要的瓶颈。传统的下一个token预测仅监督上下文之后的内容,却将该延续部分背后的中间推理过程隐含起来。我们提出了REER-PT,这是一个可扩展的框架,将逆向工程推理(Reverse-Engineered Reasoning,REER)扩展到原始预训练数据中。REER-PT能够识别出那些难以预测但仍可从先前上下文中推断出的延续部分,并插入简洁的推理注释,以重建上下文与延续部分之间缺失的联系。候选注释在离线状态下生成并优化,其中困惑度(perplexity)被用作优化信号。长度约束和目标泄漏(target leakage)会过滤掉无用或琐碎的注释。这种稀疏转换保留了源文本,并且与标准的下一个token预测兼容,避免了预训练期间的在线推理推出。我们应用REER-PT将源预训练语料库转换为增强后的语料库。在增强数据、原始token和选定延续部分的比较中,困惑度降低幅度为0.42至7.29,且注释的13-grams中仅有约0.05%逐字出现在源文本中。随后,我们使用相同的架构和训练配置,分别在源语料库和增强语料库上训练两个具有6.8亿参数的模型。增强数据训练的模型在多个知识和推理基准上的性能提升高达2.07个百分点。困惑度分析表明延续部分的可预测性得到提升,而受控预训练实验表明,这种数据增强可以在不改变标准预训练目标的情况下提高模型性能。
英文摘要
As language-model compute continues to scale, high-quality training data is becoming an increasingly important bottleneck. Conventional next-token prediction supervises what follows a context but leaves the intermediate reasoning behind that continuation implicit. We introduce \textbf{REER-PT}, a scalable framework that extends Reverse-Engineered Reasoning (REER) to raw pre-training data. REER-PT identifies continuations that are difficult to predict but can still be inferred from the preceding context, and inserts concise reasoning annotations that reconstruct the missing connection between context and continuation. Candidate annotations are generated and refined offline, with perplexity serving as the optimization signal. Constraints on length and target leakage filter out unhelpful or trivial annotations. This sparse transformation preserves the source text and remains compatible with standard next-token prediction, avoiding online reasoning rollouts during pre-training. We apply REER-PT to transform a source pre-training corpus into an augmented one. Across augmented-data, original-token, and selected-continuation comparisons, perplexity reductions range from 0.42 to 7.29, and only about 0.05\% of annotation 13-grams appear verbatim in the source text. We then train two 680M-parameter models with the same architecture and training configuration on the source and augmented corpora, respectively. The augmented-data model gains up to 2.07 percentage points on several knowledge and reasoning benchmarks. Together, the perplexity analysis indicates improved continuation predictability, while the controlled pre-training experiments suggest that this augmentation can improve model performance without changing the standard pre-training objective.