arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RLVR中的验证器错误:奖励黑客、反馈极限与选择性控制

Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

Christian Moya, Elliott Thornley, Guang Lin

arXiv 2609.35677首次发表:更新:

发表机构

Purdue University; National University of Singapore(普渡大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RLVR中不完美验证器导致的奖励黑客问题,本文利用梯度流刻画其发生条件,证明观测数据不足以检测和纠正错误,并提出基于审计反馈的选择性控制方法,实验验证其能减少被接受的错误并提升正确性。

AI 中文摘要

在可验证奖励的强化学习(RLVR)中,不完美的验证器可能奖励不正确的响应,从而为奖励黑客行为创造机会。利用固定验证器的梯度流,我们刻画了奖励上升而正确性下降的条件。随后,我们证明RLVR过程中可观测到的信息通常不足以检测或识别被接受的错误,也不足以在保证不牺牲正确响应的前提下确保这些错误的减少。为应对这一局限,我们利用来自审计的额外正确性反馈构建了一种修正机制。该修正实现了“选择性控制”:在当前策略下,只要该修正能抵消验证器奖励对错误行为的压力,它就能降低被接受错误的概率并提高正确响应的概率。基于对数线性模型、神经上下文赌博机以及语言模型的实验支持了上述分析,并表明在部分审计下的选择性控制能够在提高正确性的同时减少被接受的错误。

英文摘要

In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑