发表机构
Jilin University; Singapore University of Technology and Design(吉林大学; 新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过交叉实验和受控干预,发现RLVR训练会使模型固化学得的报告惯例,与当前请求冲突,降低请求遵从度,且惯例匹配准确率不足以全面评估训练后行为。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)已成为一种通过自动检查答案来提升语言模型在推理任务上性能的重要方法。然而,与惯例匹配的评估无法揭示强化一种报告惯例是否会降低模型对初始策略已遵循的另一请求的遵从度。为测试这一点,我们在两种报告惯例下训练匹配的策略,并在两种当前请求下分别评估每个策略,使用相同的初始策略作为共享参考。我们通过受控干预和独立人工校准补充了这一交叉设计。在GSM8K上,相对于95.45%的初始基线,框格式RLVR在五个训练种子中的四个中将Qwen2.5-7B响应中包含所请求的哈希格式有效载荷的比例降低了35.33至74.37个百分点;第五个种子提高了2.50个百分点。在四个性能下降的运行中,几乎所有省略所请求有效载荷的响应都保留了训练所得的框式惯例,且相同的四个种子在两种固定改写下均出现性能下降。仅改变监督目标中的最终答案标记,就逆转了模型在三个种子中的报告惯例偏好,这为这种偏好的可学习性提供了受控证据。在三个独立人工校准的设置中,惯例敏感评分器下的增益超过了相应的承诺答案正确性(即模型实际承诺答案的正确性)增益。综合来看,这些结果区分了三种不同的训练后结果:学得的报告偏好、当前请求遵从度和承诺答案正确性。它们表明,仅凭惯例匹配的准确率并不能完全表征训练后行为,并促使在评估惯例匹配的任务准确率的同时,也评估当前请求遵从度。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using automatically checked answers. Yet convention-matched evaluation cannot reveal whether reinforcing one reporting convention reduces adherence to a different request that the initial policy already follows. To test this, we train matched policies under two reporting conventions and evaluate each policy under both current requests, using the same initial policy as a shared reference. We complement this crossed design with controlled interventions and independent human calibration. On GSM8K, boxed-format RLVR reduces the fraction of Qwen2.5-7B responses containing the requested hash-format payload by 35.33--74.37 percentage points relative to a 95.45% initial baseline in four of five training seeds; the fifth improves by 2.50 points. In the four deteriorating runs, almost every response that omits the requested payload instead retains the trained boxed convention, and the same four seeds deteriorate under two fixed paraphrases. Changing only the final-answer marker in supervised targets reverses which reporting convention the model prefers across three seeds, providing controlled evidence that this preference is learnable. Across three settings with independent human calibration, gains under a convention-sensitive scorer exceed the corresponding gains in committed-answer correctness, i.e., the correctness of the answer the model actually commits to. Together, these results separate three distinct post-training outcomes: learned reporting preference, current-request adherence, and committed-answer correctness. They show that convention-matched accuracy alone does not fully characterize post-training behavior and motivate evaluating current-request adherence alongside convention-matched task accuracy.