发表机构
University of Science and Technology of China; Shanghai University of Finance and Economics; University of Cambridge; Nanyang Technological University; EPIC Lab, Shanghai Jiao Tong University; Fudan University; University of Electronic Science and Technology of China; Imperial College London(中国科学技术大学; 上海财经大学; 剑桥大学; 南洋理工大学; 上海交通大学EPIC实验室; 复旦大学; 电子科技大学; 伦敦帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出任务-工具协同进化框架,通过在线和任务后自我改进递归优化推理数据合成工具,生成更难任务,提升下游模型性能。
AI 中文摘要
生成逐渐变难的推理问题需要随任务分布演变而适应的合成程序。现有的任务级递归将生成的问题作为种子重用,但保持构建工具不变。我们提出任务-工具协同进化,一种用于推理数据合成中递归工具自我改进(RSI)的框架。在线自我改进在生成过程中将中间求解器失败转化为可重用技能。任务后自我改进在每批之后修订技能、提示和工作流程,仅在候选在有限成本增加内生成更难的有效任务时采用。模型权重和验证标准保持不变。在数学、编码和科学领域,平均求解器准确率在十四轮进化中从100.0%降至54.8%。消融实验表明,结合两种更新计划比固定工具递归或任一单独计划产生更难的任务。所得数据改善了下游SFT和GRPO性能。特别是,一个27B学生模型在10K合成数学示例上微调,在APEX上达到62.5%的平均16准确率,与选定的前沿模型参考相当。这些结果支持将合成工具与任务一起调整,以生成具有下游训练价值的越来越具挑战性的数据。
英文摘要
Generating progressively harder reasoning problems requires synthesis procedures that adapt as the task distribution evolves. Existing task-level recursion reuses generated problems as seeds but leaves the construction harness unchanged. We present task-harness co-evolution, a framework for recursive harness self-improvement (RSI) in reasoning-data synthesis. Online self-improvement converts intermediate solver failures into reusable skills during generation. Post-task self-improvement revises skills, prompts, and workflows after each batch, adopting candidates only when they generate harder valid tasks within a bounded cost increase. Model weights and verification criteria remain fixed. Across mathematics, coding, and science, mean solver accuracy decreases from 100.0% to 54.8% over fourteen evolution rounds. Ablations show that combining both update schedules produces harder tasks than fixed-harness recursion or either schedule alone. The resulting data improves downstream SFT and GRPO performance. In particular, a 27B student fine-tuned on 10K synthesized mathematics examples achieves 62.5% mean-16 accuracy on APEX, competitive with selected frontier-model references. These results support adapting the synthesis harness alongside the tasks to generate increasingly challenging data with downstream training value.