arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TIPCODER:强化学习增强的代码生成测试时指令提议器

TIPCODER: Reinforcement Learning Boosted Test-time Instruction Proposer for Code Generation

Minyu Chen, Sihao Wu, Ling-I Wu, Song Qin, Jingyang Li, Lei Ning, Jianxin Xue, Guoqiang Li

arXiv 2609.03309首次发表:更新:

发表机构

Shenzhen Technology University; Institute of AI for Industries, Chinese Academy of Sciences; Shanghai Jiao Tong University; Shanghai Polytechnic University(深圳职业技术大学; 中国科学院产业智能研究院; 上海交通大学; 上海工程技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出TipCoder,一种强化学习增强的代码生成测试时指令提议器,通过生成问题特定辅助提示并结合探索-选择设计,在代码生成基准和Code LLMs上表现优于随机采样与通用提示优化基线。

AI 中文摘要

代码生成的测试时缩放通常通过从固定指令中采样多个程序来探索解空间,本文研究了一个互补方向:实例级指令空间探索。我们观察到,许多编码失败源于原始提示缺失约束、被忽视的边界情况或误导性推理路径。为解决该问题,我们提出TipCoder,一种在代码合成前生成问题特定辅助提示的测试时指令提议器。TipCoder将多轮调试轨迹提炼为主动指导,并使用边际效用奖励通过强化学习进一步优化提议器。推理时,它同时生成基础解和提示引导解,并使用奖励模型进行事后选择。这种探索-选择设计使提示能挖掘额外候选潜力,同时减少不必要指导带来的退化。在评估的代码生成基准和目标代码大语言模型(Code LLMs)上,TipCoder提供了一致的指令级测试时缩放策略,在共享的基于奖励模型的选择协议下,与随机采样和通用提示优化基线相比表现更优。

英文摘要

Test-time scaling for code generation typically explores the solution space by sampling multiple programs from a fixed instruction. We study a complementary direction: instance-level instruction-space exploration. Our observation is that many coding failures stem from missing constraints, overlooked edge cases, or misleading reasoning paths induced by the original prompt. To address this, we propose TipCoder, a test-time instruction proposer that generates problem-specific auxiliary tips before code synthesis. TipCoder distills multi-turn debugging trajectories into proactive guidance and further optimizes the Proposer with reinforcement learning using a marginal-utility reward. At inference time, it generates both a base solution and a tip-guided solution, and applies a Reward Model for post-hoc selection. This exploration-selection design allows tips to expose additional candidate potential while reducing regressions from unnecessary guidance. Across the evaluated code-generation benchmarks and target Code LLMs, TipCoder provides a consistent instruction-level test-time scaling strategy, comparing favorably with stochastic sampling and generic prompt optimization baselines under a shared reward-model-based selection protocol.

Comments15 pages, accepted by findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑