arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

递归工具自我改进用于前沿推理数据合成

Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis

Wenlong Zhang, Zhengbo Jiao, Chenxu Zhang, Lekang Jiang, SiYuan Ma, Qituan Zhang, Guo Chen, Linfeng Zhang

arXiv 2610.03548首次发表:更新:

发表机构

University of Science and Technology of China; Shanghai University of Finance and Economics; University of Cambridge; Nanyang Technological University; EPIC Lab, Shanghai Jiao Tong University; Fudan University; University of Electronic Science and Technology of China; Imperial College London(中国科学技术大学; 上海财经大学; 剑桥大学; 南洋理工大学; 上海交通大学EPIC实验室; 复旦大学; 电子科技大学; 伦敦帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出任务-工具协同进化框架,通过在线和任务后自我改进递归优化推理数据合成工具,生成更难任务,提升下游模型性能。

AI 中文摘要

生成逐渐变难的推理问题需要随任务分布演变而适应的合成程序。现有的任务级递归将生成的问题作为种子重用,但保持构建工具不变。我们提出任务-工具协同进化,一种用于推理数据合成中递归工具自我改进(RSI)的框架。在线自我改进在生成过程中将中间求解器失败转化为可重用技能。任务后自我改进在每批之后修订技能、提示和工作流程,仅在候选在有限成本增加内生成更难的有效任务时采用。模型权重和验证标准保持不变。在数学、编码和科学领域,平均求解器准确率在十四轮进化中从100.0%降至54.8%。消融实验表明,结合两种更新计划比固定工具递归或任一单独计划产生更难的任务。所得数据改善了下游SFT和GRPO性能。特别是,一个27B学生模型在10K合成数学示例上微调,在APEX上达到62.5%的平均16准确率,与选定的前沿模型参考相当。这些结果支持将合成工具与任务一起调整,以生成具有下游训练价值的越来越具挑战性的数据。

英文摘要

Generating progressively harder reasoning problems requires synthesis procedures that adapt as the task distribution evolves. Existing task-level recursion reuses generated problems as seeds but leaves the construction harness unchanged. We present task-harness co-evolution, a framework for recursive harness self-improvement (RSI) in reasoning-data synthesis. Online self-improvement converts intermediate solver failures into reusable skills during generation. Post-task self-improvement revises skills, prompts, and workflows after each batch, adopting candidates only when they generate harder valid tasks within a bounded cost increase. Model weights and verification criteria remain fixed. Across mathematics, coding, and science, mean solver accuracy decreases from 100.0% to 54.8% over fourteen evolution rounds. Ablations show that combining both update schedules produces harder tasks than fixed-harness recursion or either schedule alone. The resulting data improves downstream SFT and GRPO performance. In particular, a 27B student fine-tuned on 10K synthesized mathematics examples achieves 62.5% mean-16 accuracy on APEX, competitive with selected frontier-model references. These results support adapting the synthesis harness alongside the tasks to generate increasingly challenging data with downstream training value.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑