发表机构
Institute of Automation, Chinese Academy of Sciences; Zhongguancun Academy(中国科学院自动化研究所; 中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模强化学习后训练实验发现,逐问题增益在独立运行间高度共享,但先验信号预测力弱,而现有检查点能更准确预测新运行的改进,并提议噪声上限作为评估标准。
AI 中文摘要
数据选择、课程设计和后训练配方的评估都假设我们能在训练前判断模型将在哪些问题上有所改进。我们直接检验了这一假设。对于两个基础模型,DeepSeek-R1-Distill-Qwen-1.5B 和 Qwen2.5-Math-1.5B,我们在多达1532个竞赛数学问题上评估了十八次后训练运行,每个问题有多个样本,并将训练前可获得的信号与由不相交运行子集间一致性得出的噪声上限进行比较。两个基础模型上均得出三项发现。后训练所修复的内容是共享的:独立运行在哪些很少被解决的问题会得到改进上达成一致,噪声上限约为0.9,但两个基础模型彼此间的一致性仅为 rho=0.25——共享成分属于基础模型,而非问题本身。先验信号仅捕获了其中少数:基础通过率、正确解的可能性、更大模型的通过率及其组合仅解释了增益中可解释方差的0.30和0.24。现有检查点是更好的预测器:来自另一家族的单个检查点的逐问题增益比每个先验信号(单独或组合)更好地预测新运行(0.52对0.33,以及0.34对0.17,相对于组合信号)。这些结论在2025-2026年竞赛的问题上以及当每次比较两侧的基线被独立估计时均成立。我们提议将噪声上限作为逐问题信号的标准伴随指标。
英文摘要
Data selection, curricula and the evaluation of post-training recipes all assume that we can tell, before training, which problems a model will improve on. We test this assumption directly. For two base models, DeepSeek-R1-Distill-Qwen-1.5B and Qwen2.5-Math-1.5B, we evaluate eighteen post-training runs on up to 1532 competition math problems with many samples per problem, and compare signals available before training against a noise ceiling derived from the agreement between disjoint subsets of runs. Three findings hold on both base models. What post-training fixes is shared: independent runs agree on which rarely solved problems improve, with a noise ceiling of about 0.9, yet the two base models agree with each other only at rho=0.25 - the shared component belongs to the base model, not to the problem. A priori signals capture a minority of it: base pass rate, the likelihood of a correct solution, a larger model's pass rate and their combinations explain only 0.30 and 0.24 of the explainable variance in gains. Existing checkpoints are the better predictor: the per-problem gains of a single checkpoint from another family predict a new run better than every a priori signal, alone or combined (0.52 vs. 0.33 and 0.34 vs. 0.17 against the combined signals). The conclusions hold on problems from 2025-2026 competitions and when the baselines on the two sides of every comparison are estimated independently. We propose the noise ceiling as a standard companion to per-problem signals.
Comments10 pages, 2 figures