AI 中文总结
该研究针对代码智能体强化学习中反馈利用不足的问题,提出TaPR框架,通过将执行反馈转换为密集每轮测试通过率奖励,在219个代码问题上提升了多轮代码生成成功率。
AI 中文摘要
多轮代码智能体依赖执行反馈修复错误程序,但标准强化学习范式主要使用单次结果奖励优化和评估策略性能。这种错位将初始代码生成与反馈驱动优化混为一谈,丢弃了中间轮次的细粒度执行信号,且无法评估策略是否真正获得自我修复能力。我们提出感知测试策略优化(Test-aware Policy Refinement, TaPR)框架,在一致的多轮交互协议下将执行反馈转换为密集的每轮测试通过率奖励。在来自LiveCodeBench的219个代码生成问题上,TaPR在六个模型上使合并的三轮成功率(Pass@3)提升2.44个百分点。在预定义的7B/8B高余量切片中,合并准确率从30.25%升至33.56%(+3.31个百分点),配对试验中有42次提升和13次下降。在匹配的Qwen3-8B ablation实验中,密集奖励在所有前十步提供非零反馈,在测试预算内达到比仅结果型GRPO更高的难子集峰值,尽管GRPO在第300步时几乎匹配合并的Pass@3。我们的主要贡献是一种奖励分解框架和轮次感知评估协议,将首次生成质量与多轮修复能力解耦。
英文摘要
Multi-turn code agents rely on execution feedback to repair incorrect programs, yet standard reinforcement learning paradigms optimize and evaluate policy performance primarily using single-shot outcome rewards. This misalignment conflates initial code generation with feedback-driven refinement, discards granular execution signals across intermediate turns, and fails to evaluate whether the policy actually acquires self-repair capabilities. We propose Test-aware Policy Refinement (TaPR), a framework that transforms execution feedback into a dense per-turn test-pass-ratio reward under a consistent multi-turn interaction protocol. Across six models on 219 code-generation problems from LiveCodeBench, TaPR improves the pooled three-turn success rate (Pass@3) by 2.44 percentage points. In the predefined 7B/8B high-headroom slice, pooled accuracy increases from 30.25% to 33.56% (+3.31 pp), with 42 improvements and 13 regressions in paired trials. On a matched Qwen3-8B ablation, the dense reward supplies nonzero feedback in all of the first ten steps and reaches a higher Hard-subset peak than outcome-only GRPO within the tested budget, although GRPO nearly matches pooled Pass@3 by step 300. Our primary contribution is a reward-decomposition framework and a turn-aware evaluation protocol that decouple first-shot generation quality from multi-turn repair competence.
Comments9 pages, 3 figures, 3 tables. Aofan Liu and Jingxiang Meng contributed equally; Fangxin Liu and Yongbiao Chen are corresponding authors