arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理核心:为完成监督式推理训练设计广谱过程化数据

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training

Damien Sileo, Valentin Lacombe, Dimitri Kachler

arXiv 2608.05148首次发表:更新:

发表机构

Univ. Lille; Inria; CNRS; Centrale Lille(里尔大学; 法国国家信息与自动化研究所; 法国国家科学研究中心; 里尔中央理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出含50类任务生成器的Reasoning Core集合,经对比实验,其在3B规模模型的三个推理基准上得分最优,相关资源已公开。

AI 中文摘要

过程化生成器可大规模生成有用且可验证的推理问题,但作为完成监督式微调的数据却较少受到关注。我们推出Reasoning Core,这是一个包含50个生成器的集合,涵盖数学、逻辑、规划、状态跟踪、形式语言、结构化数据、游戏、因果关系和代码领域,具备语义评分器、难度控制和任务评估器。在匹配的完成监督式协议下,我们在四个基础模型设置和多个训练时长中,将Reasoning Core与Procedural Warmup、Reasoning Gym和SynLogic进行对比。在主要的3B规模模型对比中,Reasoning Core在DROP、LogiQA和ARC-Challenge上取得最高平均分数,超过了无过程化数据的基线以及所有三种替代过程化集合。任务层面分析显示,仅语义有效性无法确保训练效用,凸显紧凑目标和校准难度是重要的设计因素。我们开展了结合模型辅助审查、人工裁决和回归测试的审计,应用于Reasoning Core开发全过程及其他集合,结果揭示了生成、渲染、目标和评分之间存在细微不匹配,提醒人们仅过程化生成无法保证正确性。该库、生成的数据集和审计材料均已公开可用。

英文摘要

Procedural generators produce useful verifiable reasoning problems at scale, but have received less attention as data for completion-supervised fine-tuning. We introduce Reasoning Core, a collection of 50 generators spanning mathematics, logic, planning, state tracking, formal languages, structured data, games, causality, and code, with semantic scorers, difficulty controls, and task evaluators. Under a matched completion-supervised protocol, we compare Reasoning Core with Procedural Warmup, Reasoning Gym, and SynLogic across four base-model settings and multiple training durations. In the primary 3B comparison, Reasoning Core achieves the highest mean scores on DROP, LogiQA, and ARC-Challenge, exceeding both the baseline without procedural data and all three alternative procedural collections. Task-level analyses show that semantic validity alone does not ensure training utility, highlighting compact targets and calibrated difficulty as important design factors. We ran audits combining model-assisted review, human adjudication, and regression testing. Applied throughout Reasoning Core development and to the other collections, they reveal subtle mismatches among generation, rendering, targets, and scoring, a reminder that procedural generation alone does not guarantee correctness. The library, generated datasets, and audit material are publicly available.

Comments20 pages, 3 figures. Code: https://github.com/sileod/reasoning-core Data: https://hf.co/collections/reasoning-core/datasets

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑