AI 中文总结
本文针对竞赛数学强化微调的三大困难,提出Question-begets-Question(QbQ)方法及自演进课程,使Qwen2.5-Math-7B的AIME任务pass@1从5.6%提升至16.5%,且能解决更难问题。
AI 中文摘要
让语言模型掌握未精通技能时,存在三大反复出现的困难:训练数据稀缺、通常无法获取真实推理轨迹、模型常出现明显上限,超过该上限后增加数据也无法进一步提升性能。我们在受控环境中研究这些困难,对Qwen2.5-Math-7B在竞赛数学(AIME)任务上进行微调,该模型最初仅能解决5.6%的问题(pass@1指标)。为解决数据稀缺问题,我们提出Question-begets-Question(QbQ),这是一种可扩展的流程,其中教师将现有问题转化为探究相同基础技能的多样变体;为模拟缺少权威推理的情况,我们仅通过问题陈述和最终答案进行强化学习训练,绝不使用教师的推理轨迹。然而,对这类数据的静态训练性能远未达到任务要求:真实加合成数据增强以及非课程QbQ生成的合成数据训练,分别将pass@1上限限制在12.5%和14.5%,尽管数据量大幅增加。我们的核心发现是,该上限并非模型固有属性。我们提出一种自演进课程,每轮都会评估当前检查点,从模型基本能答对的问题中生成QbQ种子,并基于得到的变体进行训练;在相同数据预算下,这打破了上限,使pass@1提升至16.5%,且经过20轮后仍无饱和迹象。与直觉相反,我们发现模型在基本能答对的问题变体上训练时性能会提升,且经此训练的模型后续能解决训练期间从未见过的更难问题。
英文摘要
Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable, and models often exhibit an apparent ceiling beyond which additional data yields no further improvement. We study these difficulties in a controlled setting, fine-tuning Qwen2.5-Math-7B on competition mathematics (AIME), a task on which it initially solves only 5.6\% of problems (pass@1). To address data scarcity, we introduce Question-begets-Question (QbQ), a scalable procedure in which a teacher transforms existing problems into diverse variants that probe the same underlying skills; to model the absence of oracle reasoning, we train exclusively via reinforcement learning on problem statements and final answers, never on teacher reasoning traces. Static training on such data, however, plateaus well short of the task: real-plus-synthetic augmentation and non-curriculum QbQ generated synthetic data training cap pass@1 at 12.5\% and 14.5\% respectively, despite large increases in data. Our central finding is that this ceiling is not intrinsic to the model. We propose a self-evolving curriculum that, each round, evaluates the current checkpoint, seeds QbQ from the problems it can mostly get right, and trains on the resulting variants; under an identical data budget, this breaks the ceiling and lifts pass@1 to 16.5\% with no sign of saturation after 20 rounds. Counterintuitively, we find that models improve when trained on variants of problems they can mostly get right, and that models trained this way go on to solve harder problems never seen during training.