arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22471cs.SEcs.AI

MineValiCoder:通过测试用例质量挖掘和基于二分图的相互验证实现可靠的代码生成

MineValiCoder: Reliable Code Generation with Test Case Quality Mining and Bipartite Graph-Based Mutual Validation

Zhen Zhao, Qihang Yang, Feifei Dai, Xiangfang Li, Bo Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对基于大语言模型的测试驱动开发中现有方法的不足,提出MineValiCoder框架,通过测试用例质量挖掘、并行TDD优化和二分图代码-测试相互验证三个模块,有效减轻LLM随机性,提升自动代码生成可靠性,在多基准测试中表现优异。

中文摘要 AI 辅助

基于大语言模型(LLM)的测试驱动开发(TDD)推动了自动代码生成。现有方法严重依赖人工编写的测试用例,仅根据自然语言需求时无法有效运行。虽近期工作能自动生成测试,但常忽视LLM的内在随机性,导致错误测试产生误导反馈,混合质量的测试用例产生冲突评估信号。为应对这些挑战,我们提出MineValiCoder,一个基于测试用例质量和代码质量相互强化的协作闭环TDD框架。它由三个模块组成:测试用例质量挖掘(TCQM)模块通过自我验证过滤错误测试用例,提供可靠优化监督;并行TDD优化模块使用经过验证的测试用例反馈迭代优化代码并生成多样高质量代码候选;基于二分图的代码-测试相互验证(BiCoTeV)模块动态建模代码-测试交互并进行相互验证评分以实现稳定可靠的最优代码选择。在四个LLM和主流基准上的广泛评估表明,MineValiCoder显著优于现有方法。具体而言,它在HumanEval上的Pass@1分数为96.34%,在MBPP上为87.40%,在APPS上为64.00%,在LiveCodeBench上为51.33%。这些结果证明了MineValiCoder在减轻LLM随机性和提高自动代码生成可靠性方面的有效性。

英文摘要

Large Language Model (LLM)-based Test-Driven Development (TDD) has advanced automated code generation. However, existing approaches depend heavily on human-crafted test cases and cannot operate effectively when only natural-language requirements are available. Although recent work enables automatic test generation, it often overlooks the inherent stochasticity of LLMs, leading to two key defects: faulty tests generate misleading feedback that distorts code optimization, while mixed-quality test cases produce conflicting evaluation signals that hinder reliable code selection. To address these challenges, we propose MineValiCoder, a collaborative closed-loop TDD framework based on the mutual reinforcement of test-case quality and code quality. MineValiCoder comprises three modules. The Test Case Quality Mining (TCQM) module filters faulty test cases through self-validation, providing reliable optimization supervision. The Parallel TDD Refinement module iteratively optimizes code and generates diverse high-quality code candidates using validated test-case feedback. The Bipartite Graph-Based Code-Test Mutual Validation (BiCoTeV) module dynamically models code-test interactions and performs mutual validation scoring for stable and reliable optimal-code selection. Extensive evaluations across four LLMs and mainstream benchmarks show that MineValiCoder significantly outperforms state-of-the-art methods. Specifically, it achieves Pass@1 scores of 96.34% on HumanEval, 87.40% on MBPP, 64.00% on APPS, and 51.33% on LiveCodeBench. These results demonstrate the effectiveness of MineValiCoder in mitigating LLM stochasticity and improving the reliability of automated code generation.

发表机构

  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
  • University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑