面向代码大语言模型的健全性与对抗性测试用例生成的两阶段强化学习
Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs
浏览论文内容
中文总结 AI 辅助
针对代码LLM测试用例稀缺问题,提出两阶段RL框架TCS,在TACO和LiveCodeBench上提升pass@1与答案选择能力,且可用于其他LLM输出选择。
中文摘要 AI 辅助
强化学习(RL)通过可执行反馈大幅推动了基于大语言模型(LLM)的代码生成。编程问题的反馈主要来自特定测试用例,而高质量测试用例往往稀缺,因为它们需同时具备健全性与判别性。因此,我们研究利用学习到的模型自动生成测试用例,发现这本质上是一个对抗性强化学习问题:模型需根据求解器当前的失败模式生成有效测试用例作为反例。我们提出了测试用例缩放(Test Cases Scaling, TCS),这是一个用于有效测试用例生成的两阶段强化学习框架。两个阶段均从滚动策略对齐缓冲区训练测试生成器:第一阶段生成与参考解一致的测试,第二阶段将缓冲区限制为当前失败模式并学习反例测试。在TACO和LiveCodeBench数据集上,TCS根据生成的测试同时提升了pass@1指标和推理时答案选择能力。我们还发现,学习到的测试生成器也能在其他LLM输出中实现有效选择。
英文摘要
Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminative. We thus turn to study the auto-generation of test cases using the learned model. We find this is naturally an adversarial RL problem: the model is expected to generate effective test cases as counterexamples, depending on the solver's current failure modes. We propose Test Cases Scaling (TCS), a two-stage RL framework for effective test generation. Both stages train a test generator from a rolling policy-aligned buffer: Stage 1 generates tests consistent with the reference solution, and Stage 2 restricts the buffer to current failure modes and learns counterexample tests. Across TACO and LiveCodeBench, TCS improves both pass@1 and inference-time answer selection according to generated tests. We find the learned test generator also enables effective selection among other LLM outputs.
发表机构
- Nanyang Technological University(南洋理工大学)
- Skywork AI(天工智源)
机构由 AI 辅助整理,请以论文原文为准。