arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21208cs.AI

基于多样性剪枝测试的信息增益奖励:GT锚定验证器协同训练实现可靠代码生成

Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation

Ana Nunez, Peyman Najafirad

首次发表
浏览论文内容

中文总结 AI 辅助

提出CoVer框架,通过信息增益奖励和多样性剪枝测试解决代码生成协同训练中的宽松性崩溃与集中性偏差,在多个基准上显著提升pass@1,并增强代码选择性能。

中文摘要 AI 辅助

将单一语言模型同时作为编码器和测试作者进行协同训练的自博弈方法,有望将代码生成强化学习从固定测试套件中解放出来,但此类方法面临两种相互关联的病理现象:一是宽松性崩溃,即通过琐碎、无区分度的测试使通过率奖励最大化;二是集中性偏差,即独立同分布采样的测试聚集在模态输入上,导致估计器方差膨胀。我们提出CoVer(协同训练的编码器与验证器),一种单策略GRPO框架,同时解决上述两种失败模式。首先,信息增益(IG)奖励根据每个自生成测试的通过/失败向量与分级、地面真值锚定的正确性信号y [0, 1] m之间的互信息为其评分,并通过二者协方差的符号进行门控,使得只有具有正向区分度的测试才能获得奖励。其次,一个三阶段多样性感知选择步骤将候选池剪枝为行为上非冗余的测试套件(无效性、输入字符串、执行配置文件过滤),在固定执行预算下提高IG估计器的有效样本量。在五个基准(LiveBench、MBPP、LiveCodeBench、CodeContests、CodeForces)上,CoVer在7B规模下将一次性pass@1提升+5.8个百分点,在14B规模下提升+7.1个百分点,相较于Qwen2.5-Instruct骨干模型,并在两个规模下均取得所有对比方法中最高的宏平均分。作为CodeT排序流程中的即插即用骨干,CoVer-7B额外增加+3.5个百分点,展示了协同训练在生成和选择两方面的双重益处。

英文摘要

Self-play methods that co-train a single language model as both coder and test author promise to move code-generation RL beyond fixed test suites, but they suffer from two coupled pathologies: permissiveness collapse, where pass-rate rewards are maximised by trivial, non-discriminative tests, and concentration bias, where i.i.d. sampled tests cluster on modal inputs and inflate estimator variance. We introduce CoVer (Co-trained Coder and Verifier), a single-policy GRPO framework that addresses both failure modes. First, an information-gain (IG) reward scores each self-generated test by the mutual information between its pass/fail vector and a graded, ground-truth-anchored correctness signal y [0, 1] m, gated by the sign of their covariance so that only positively discriminative tests receive reward. Second, a three-stage diversity-aware selection step prunes a candidate pool to a behaviourally non-redundant suite (invalidity, input-string, execution-profile filtering), raising the effective sample size of the IG estimator at fixed execution budget. On five benchmarks (LiveBench, MBPP, LiveCodeBench, CodeContests, Code-Forces), CoVer raises one-shot pass@1 by +5.8 points at 7B and +7.1 points at 14B over the Qwen2.5-Instruct backbone, and achieves the highest macro-average among all compared methods at both scales. As a drop-in backbone inside the CodeT ranking pipeline, CoVer-7B adds +3.5 points, demonstrating the dual benefit of co-training for both generation and selection.

发表机构

  • Secure AI and Autonomy Lab(安全人工智能与自主实验室)
  • University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

机构由 AI 辅助整理,请以论文原文为准。

↑