arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06041cs.SE

SpecCoder:基于规范感知的代码生成与课程式双任务强化学习

SpecCoder: Specification-Aware Code Generation with Curriculum Dual-Task Reinforcement Learning

  • School of Computer Science, Fudan University(复旦大学计算机科学技术学院)
  • Department of Information Engineering, The Chinese University of Hong Kong(香港中文大学信息工程系)

机构由 AI 辅助整理,请以论文原文为准。

Yixuan Li, Mingxuan Huang, Jiajing Wang, Weidong Yang, Xinyi Liu, Ben Fei, Lipeng Ma

AI总结:

针对大语言模型代码生成中忽略需求细节的问题,提出规范感知两阶段训练框架SpecCoder,结合规范引导SFT与课程式双任务GRPO,在多个基准上提升代码生成与智能体工作流性能。

AI中文摘要:

大语言模型(LLMs)在代码生成方面取得了显著进展,但在处理需要理解丰富自然语言需求的挑战性编程任务时仍存在困难。这些需求通常指定问题目标、输入/输出格式、约束条件、示例和边界情况。即使忽略其中一项,也可能生成可执行但功能不正确的代码。现有的免训练方法主要依赖提示或基于智能体的工作流,而基于训练的方法通常优化最终代码输出。然而,现有方法对学习从原始需求到结构化规范的中间映射以及将其落实到具体实现行为方面提供的监督有限。因此,模型可能遗漏关键约束,即使生成了明确的规范,实现也可能无法始终如一地反映该规范。受这一差距的启发,我们提出了SpecCoder,一种用于代码生成的规范感知两阶段训练框架。SpecCoder首先采用规范引导的SFT,训练LLMs推导结构化规范分析并基于这些分析生成代码。随后引入课程式双任务GRPO,联合优化规范引导的生成与判别,以增强规范与代码行为之间的对应关系。在APPS、CodeContests和xCodeEval上的实验证明了规范感知训练的有效性,SpecCoder持续提升了独立代码生成和基于智能体的工作流。在BigCodeBench-Hard和ClassEval上的额外评估,以及人工评估和扰动研究,进一步验证了结构化规范在引导代码生成与判别中的作用。

英文摘要:

Large language models (LLMs) have made substantial progress in code generation but still struggle with challenging programming tasks that require understanding rich natural language requirements. These requirements often specify problem goals, input/output formats, constraints, examples, and edge cases. Overlooking even one may produce executable but functionally incorrect code. Existing training-free methods mainly rely on prompting or agent-based workflows, while training-based methods typically optimize final code outputs. However, existing approaches provide limited supervision for learning the intermediate mapping from raw requirements to structured specifications and for grounding them in concrete implementation behavior. Consequently, models may omit critical constraints, and even when an explicit specification is produced, the implementation may fail to reflect it consistently. Motivated by this gap, we propose SpecCoder, a specification-aware two-stage training framework for code generation. SpecCoder first employs specification-guided SFT to train LLMs to derive structured specification analyses and generate code conditioned on them. It then introduces curriculum dual-task GRPO, which jointly optimizes specification-guided generation and discrimination to encourage stronger correspondence between specifications and code behavior. Experiments on APPS, CodeContests, and xCodeEval demonstrate the effectiveness of specification-aware training, with SpecCoder consistently improving both standalone code generation and agent-based workflows. Additional evaluations on BigCodeBench-Hard and ClassEval, alongside human evaluation and perturbation studies, further validate the role of structured specifications in guiding code generation and discrimination.

补充信息

↑