arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测量与缓解RLVR中的解模式崩溃

Measuring and Mitigating Solution Mode Collapse in RLVR

Liv G. d'Aliberti, Marwa Abdulhai, Sofiia Druchyna, Peter Henderson, Manoel Horta Ribeiro

arXiv 2610.11064首次发表:更新:

发表机构

Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对RLVR存在的解模式崩溃问题,构建多解任务基准ModeBench进行测量,提出Re:Max方法缓解该问题,可提升策略成功频率与成功方式数量。

AI 中文摘要

语言模型(LM)通常可以用多种方式回答同一个问题,但带可验证奖励的强化学习(RLVR)对模型产生的正确答案并不挑剔。无论解是熟悉答案的第1000个副本还是模型从未生成过的解,都会获得相同的奖励。然而,让模型在训练过程中保留多个正确解具有潜在价值,例如,多种模式可为用户提供选择,并提供提升整体模型性能的问题解决策略。本文中,我们引入ModeBench,这是一个多解任务基准,其中验证器会返回正确性和已发现的模式。随后,我们使用ModeBench测量RLVR后训练下解多样性的变化。我们发现,即使准确率保持或提升,RLVR后训练也会将概率集中到更少的正确模式上,而且前沿模型已经高度集中。接着,我们提出解决方案Re:Max,它会在回放缓冲区中存储每个已发现模式的一个已验证示例,并对这些存储的模式进行均匀训练。因此,仅发现一次的解会被练习的频率与重复发现的解相同。在三种模型规模、两种RL目标和更难的任务构造下,回放既提升了策略成功的频率,也提升了策略成功的方式数量。

英文摘要

A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produced before. Yet, there is potential value in having the model retain multiple correct solutions as it is trained. For instance, multiple modes may give users a choice and provide problem-solving strategies that improve overall model performance. Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered. We then use ModeBench to measure how solution diversity changes under RLVR post-training. We find that RLVR post-training concentrates probability onto fewer correct modes even as accuracy holds or improves, and moreover, that frontier models are already highly concentrated. We then introduce our solution, Re:Max, which stores one verified example per discovered mode in a replay buffer and trains on those stored modes uniformly. A solution found once is, therefore, practiced as often as one found repeatedly. Across three model scales, two RL objectives, and harder task constructions, replay improves both how often a policy succeeds and how many different ways it can succeed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑