发表机构
Duke University; Adobe Inc.; Oregon State University; Pennsylvania State University; National University of Singapore; Amazon(杜克大学; 奥多比公司; 俄勒冈州立大学; 宾夕法尼亚州立大学; 新加坡国立大学; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对开放式LLM自我改进中奖励验证问题,提出基于任务转换的RLSVR方法,通过SpyRL实例化,在文本摘要等任务实验中,该方法在不可验证任务上优于现有方法,在可验证推理任务上有提升,扩展了基于RLVR的自我改进。
AI 中文摘要
可验证奖励强化学习(RLVR)通过大规模优化推动了面向推理的大型语言模型(LLMs)的进展,但其适用性主要限于数学和编码等可确定性验证正确性的领域。开放式任务通常依赖人类偏好、奖励模型或基于LLM的评判,存在评估偏差等问题。基于自监督学习原理,提出了自验证奖励强化学习(RLSVR),一种基于任务转换的训练范式,将开放式任务转换为可验证的代理环境以自动生成奖励信号。通过SpyRL实例化RLSVR,实验表明SpyRL在不可验证任务上优于现有自我改进方法,在可验证推理任务上也有持续提升,证明任务转换可扩展基于RLVR的可扩展自我改进。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verifiable. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference costs. Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a Self-PlaY Reinforcement Learning method inspired by social deduction game Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains. Models and code have been released at https://github.com/wangqinsi1/RLSVR/tree/SpyRL.
CommentsCOLM 2026