发表机构
HKUST(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出RLPF方法,将执行结果转化为分阶段奖励,微调Qwen3-32B后提升了代码生成的正确性与效率,模型性能可适度迁移,验证了代码智能体可优化生成程序。
AI 中文摘要
代码模型越来越多地借助执行反馈进行训练,但多数训练信号仍仅停留在正确性层面,这在系统代码领域留下了重要空白:两个程序可通过相同测试,却在运行时效率上存在巨大差异。本研究探讨如何训练代码智能体,使其偏好更快的正确实现,而非仅将效率作为评估指标。核心难点在于运行时是一种脆弱的奖励信号:仅在程序正确时才有意义、随任务变化,且当多数采样程序无法编译或运行时几乎无法提供指导。我们提出**RLPF(基于性能反馈的强化学习)**,将执行结果转化为分阶段奖励:失败程序按执行进度排序,正确程序则按其相对于基线向专家参考的相对改进程度排序。该方法在程序正确前提供有用反馈,在正确后提供对效率敏感的反馈。使用RLPF在PerfCodeBench上微调Qwen3-32B,使正确且可运行的解决方案比例从11.1%提升至54.6%,相对效率从8.1%提升至38.6%。训练后的模型可与更强的开放权重系统竞争,其优化行为可适度迁移至EffiBench-X。进一步研究表明,模型生成的参考提供有用但较弱的监督,且完整的复合奖励比仅正确性或仅运行时的基线更可靠。这些结果表明,代码智能体不仅可被训练为通过测试,还可优化其生成的程序。
英文摘要
Code models are increasingly trained with execution feedback, but most training signals still stop at correctness. This leaves an important gap for systems code: two programs can pass the same tests while differing greatly in runtime. We study how to train code agents to prefer faster correct implementations, rather than treating efficiency only as an evaluation metric. The key difficulty is that runtime is a fragile reward. It is meaningful only after a program is correct, varies across tasks, and gives little guidance when most sampled programs fail to compile or run. We propose \textbf{RLPF}, reinforcement learning from performance feedback, which turns execution outcomes into a staged reward. Failed programs are ordered by execution progress, while correct programs are ranked by their relative improvement from the baseline toward the expert reference. This gives useful feedback before correctness and performance-sensitive feedback after correctness. Fine-tuning Qwen3-32B with RLPF on PerfCodeBench raises correct-and-runnable solutions from $11.1\%$ to $54.6\%$ and improves relative efficiency from $8.1\%$ to $38.6\%$. The trained model becomes competitive with stronger open-weight systems, and its optimization behavior transfers modestly to EffiBench-X. Additional studies show that model-generated references provide useful but weaker supervision, and that the full composite reward is more reliable than correctness-only or runtime-only baselines. These results suggest that code agents can be trained not only to pass tests, but also to optimize the programs they write.