参与奖探针:随机奖励强化学习作为大语言模型能力的探针
Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
- University of Toronto(多伦多大学)
- Stevens Institute of Technology(史蒂文斯理工学院)
- University of California, Santa Cruz(加州大学圣克鲁兹分校)
- University of Waterloo(滑铁卢大学)
- Vector Institute(向量研究所)
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出将随机奖励强化学习作为探针工具,通过虚假奖励悖论探测大语言模型的可达性,发现训练响应存在三个不同阶段,并解决了探针性能揭示能力还是学习任务的标签泄漏问题。
AI中文摘要:
我们将虚假奖励悖论与模型的可达性联系起来,并提出随机奖励强化学习(RL)作为探针研究中的一个有用工具,从而解决了一个长达十年的争论:探针性能究竟揭示了模型的什么。关于即使随机奖励也能提升大语言模型(LLM)性能这一令人惊讶的发现,目前有两种主流解释:一种将其归因于RL训练中的特定机制;另一种则归因于数据污染。我们的结果支持一种不同的观点:虚假奖励RL可以探测模型的可达性,即在指定约束下,从当前状态进一步训练所能达到的水平,这超出了其当前性能所反映的范围。例如,两个在合成算术任务上准确率相同(3.5%)的OLMo检查点,在相同的正确性奖励RL下,其最佳运行分别达到了8.5%和55%。对OLMo检查点在预训练和中途训练阶段的考察揭示了三种不同的训练响应机制:早期,即使正确答案获得奖励,RL带来的改进也很小;预训练后期,奖励正确答案变得有效,而随机奖励仍然作用微弱;进入中途训练后,即使是随机奖励也能带来大幅提升。类似的排序也出现在对这些检查点的数字掩码监督微调(SFT)分析中,表明该模式并非特定于某种RL机制。此外,随机奖励RL为训练能使LLM做什么提供了独特的视角,因为其奖励信号不提供任何关于哪些答案正确的信息。通过询问在没有正确性反馈的情况下训练能达到什么效果,它解决了基于可解码性探针中一个核心问题的标签泄漏方面:成功的探针是揭示了模型的能力,还是学习了任务本身。
英文摘要:
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for the surprising finding that even random rewards can improve the performance of large language models (LLMs): one attributes the gains to particular mechanisms within RL training; the other to data contamination. Our results motivate a different view: spurious-reward RL can probe a model's reachability, or what further training can attain from its current state under specified constraints, beyond what is reflected in its current performance. Two OLMo checkpoints with the same accuracy on synthetic arithmetic (3.5%), for example, reach 8.5% and 55% in their best runs under the same correctness-rewarded RL. Examining OLMo checkpoints across pre-training and mid-training reveals three distinct regimes of training response: early on, RL produces little improvement even when correct answers are rewarded; later in pre-training, rewarding correct answers becomes effective while random rewards remain weak; and, upon entering mid-training, even random rewards can produce large gains. A similar ordering appears in a number-masked supervised fine-tuning (SFT) analysis of these checkpoints, suggesting that the pattern is not specific to a particular RL mechanism. Moreover, RL with random rewards offers a distinctive perspective on what training can make an LLM do, since its reward signal supplies no information about which answers are correct. By asking what training can attain without correctness feedback, it addresses the label-leakage side of a central problem in decodability-based probing: whether a successful probe reveals the model's capabilities or learns the task itself.