发表机构
Leiden Institute of Advanced Computer Science, Leiden University(莱顿大学高级计算机科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在小型模型上验证了可验证奖励强化学习(RLVR)结合搜索的可行性,发现稀疏精确匹配奖励不适用,需针对小型模型设计专门奖励。
AI 中文摘要
可验证奖励的强化学习(RLVR)在具有明确奖励的问题(如数学和编码)上表现良好,但在奖励不太明确的情况下是否同样有效仍是一个开放问题。reason-over-search 方法将 RLVR 应用于开放域问答,其中检索为答案提供依据,与参考答案的匹配提供奖励。迄今为止,该方法已在大型模型上得到验证,而在十亿参数以下的模型中,仅能通过从较大教师模型进行蒸馏来实现。我们在一个小型模型上测试了该方法。我们使用组相对策略优化(GRPO)和交错式维基百科搜索工具在 MuSiQue 上训练 Qwen3.5-0.8B,仅改变三种奖励形状,每种形状使用三个随机种子,并在七个基准问答套件的留出集上评估每个检查点。该方法有效:最佳运行的平均精确匹配达到 0.352,而未训练基线为 0.092,提升了 3.8 倍,且训练循环中没有蒸馏步骤。奖励形状也很重要。在匹配的训练时间范围内,Search-R1-faithful 的仅精确匹配奖励在每种种子下都是三种奖励中最差的,并且即使在精确匹配(它直接优化的指标)上也是最差的。我们得出结论,稀疏精确匹配奖励(RLVR 在数学和代码中的默认设置)对于这种规模的模型是错误的起点。reason-over-search 设置为小型模型上的 RLVR 提供了合适的奖励,但小型模型的 RLVR 需要其自身的奖励设计研究,而不是大型模型方法的缩小版副本。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher. We test the recipe on a small model. We train Qwen3.5-0.8B with Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia-search tool on MuSiQue, varying only the reward across three shapes over three seeds each, and we evaluate every checkpoint held-out on a seven-benchmark question-answering suite. The recipe works: the best run reaches 0.352 average exact match against a 0.092 untrained floor, a 3.8-fold gain, with no distillation step in the training loop. The reward shape also matters. The Search-R1-faithful exact-match-only reward is the worst of the three at every seed at the matched training horizon, and it is worst even on exact match, the metric it directly optimises. We conclude that the sparse exact-match reward, RLVR's default in mathematics and code, is the wrong starting point for models of this size. The reason-over-search setting can supply a suitable reward for RLVR on small models, but small-model RLVR needs its own reward-design study rather than a scaled-down copy of a large-model recipe.