发表机构
The Pennsylvania State University; Honda Research Institute USA; Honda Research Institute Japan(宾夕法尼亚州立大学; 美国本田研究所; 日本本田研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对推理RL中参考引导量的平衡问题,提出自适应参考引导(ARG)方法,将其应用于GRPO的全失败组,在Qwen3模型的五个数学推理基准上取得最优聚合pass@12性能。
AI 中文摘要
经过验证的参考解为推理模型的训练提供了正确的轨迹。或者,参考的前缀可用于引导模型生成自身的轨迹。我们应提供多少参考引导?我们通过前缀延续来研究这一问题:模型从参考前缀开始延续,若生成的轨迹通过验证则保留,否则回退到参考。由于两种过程均产生正确轨迹,我们将它们的分布与理想分布(即模型在成功验证条件下的自身分布)进行比较。对于一次延续,我们推导了闭式形式的KL散度,其在有界项范围内,随生成不同正确轨迹的概率与参考惊讶度(给定前缀下参考后缀的负对数概率)的乘积而减小。由于更长的前缀往往会提高前者概率但降低后者,仅延续成功与否无法确定最优引导量。基于此分析,我们从延续结果中学习跨训练问题共享的前缀选择器,无需估计成功概率或额外生成。由此得到的自适应参考引导(Adaptive Reference Guidance, ARG)在固定生成预算内构建正确轨迹,并将其应用于组相对策略优化(Group Relative Policy Optimization, GRPO)中的全失败组。在Qwen3-4B和Qwen3-8B上针对五个数学推理基准的实验表明,ARG在评估方法中实现了最高的聚合pass@12,且具有竞争力的平均采样准确率。
英文摘要
A verified reference solution provides a correct trajectory for training a reasoning model. Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own. How much reference guidance should we provide? We study this question through prefix continuation, where the model continues from a reference prefix and keeps the resulting trajectory if it passes verification, falling back to the reference otherwise. Since both procedures produce correct trajectories, we compare their distributions with the ideal distribution, the model's own distribution conditioned on successful verification. For one continuation, we derive the KL divergence in closed form, which, up to a bounded term, decreases with the product of the probability of generating a different correct trajectory and the reference surprisal, the negative log probability of the reference suffix given the prefix. Since a longer prefix tends to raise the former but lowers the latter, continuation success alone does not determine the preferred amount of guidance. From this analysis, we learn a prefix selector shared across training questions from continuation outcomes, without estimating success probabilities or additional generation. The resulting Adaptive Reference Guidance (ARG) constructs correct trajectories within a fixed generation budget, and we apply it to all-failure groups in Group Relative Policy Optimization (GRPO). Experiments on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks show that ARG achieves the highest aggregate pass@12 among the evaluated methods with competitive average sampled accuracy.