发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对代码生成中测试时强化学习因程序不可比较而失效的问题,提出探针驱动TTRL与熵正则化排名掩码策略优化(ERPO),通过行为共识奖励和保守更新显著提升编码基准的pass@1和pass@k。
AI 中文摘要
现有的测试时强化学习(TTRL)方法从具有规范答案的未标记测试时任务上的答案级自我投票中获取奖励,但这在代码生成中失效,因为程序无法通过表面形式进行比较,因此不能直接提供可用的训练信号。为了使TTRL适用于代码生成,我们提出了探针驱动的TTRL,该方法从问题陈述中构造无输出的探针输入,在这些探针上执行候选程序,并根据结果行为一致性定义探针共识奖励(PCR)。PCR为开放词汇程序提供了行为训练信号,但它并非完全可靠的验证器,并且容易因虚假共识而遭受奖励黑客攻击。因此,我们引入了熵正则化排名掩码策略优化(ERPO),该方法通过排名掩码将低PCR转化为保守的负更新,并通过熵上限控制策略漂移。在编码基准测试中,ERPO在域内适应和零样本迁移中显著提高了pass@1和pass@k。
英文摘要
Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal. To make TTRL applicable to code generation, we propose probe-driven TTRL, which constructs output-free probe inputs from the problem statement, executes candidate programs on these probes, and defines a Probe Consensus Reward (PCR) from the resulting behavioral agreement. PCR provides a behavioral training signal for open-vocabulary programs, but it is not a fully reliable verifier and remains susceptible to reward hacking through spurious consensus. We therefore introduce Entropy-Regularized Rank-Masked Policy Optimization (ERPO), which converts low PCR into conservative negative updates through rank masking and controls policy drift with an entropy ceiling. On coding benchmarks, ERPO substantially improves pass@1 and pass@k in both in-domain adaptation and zero-shot transfer.
CommentsAccepted to EMNLP 2026 Main Conference. 15 pages, 4 figures, 11 tables