Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
立场:具有可验证奖励的强化学习的隐藏成本与测量缺口
机构 * Stanford University(斯坦福大学) ; UC Berkeley(加州大学伯克利分校) ; The University of Tokyo(东京大学) ; RIKEN AIP(理化学研究所AIP) ; Waseda University(早稻田大学) ; Georgia Tech(佐治亚理工学院) ; Northwestern University(西北大学) ; UCLA(加州大学洛杉矶分校) ; UNC Chapel Hill(北卡罗来纳大学教堂山分校) ; Yale University(耶鲁大学) ; University of Waterloo(滑铁卢大学) ; Independent Researcher(独立研究者) ; CUHK(香港中文大学) ; UT Southwestern Medical Center(西南医学中心) ; National University of Singapore(新加坡国立大学) ; UIUC(伊利诺伊大学厄巴纳-香槟分校) ; Amazon AWS AI(亚马逊AWS人工智能)
AI总结 本文指出,具有可验证奖励的强化学习(RLVR)在提升大语言模型性能时,常因预算不匹配、尝试膨胀和基准数据污染等混淆因素导致收益被高估,并提出了预算匹配饱和曲线、校准跟踪、法官鲁棒性测试和污染筛查等最低标准。