arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16742cs.SEcs.AI

TDD-Agent:用于代码生成的测试驱动推理

TDD-Agent: Test-Driven Reasoning for Code Generation

Hongyue Yu, Kefan Li, Jiakun Li, Hongzheng Chai, Yuan Yuan, Rui He, Junyi Wei

首次发表
浏览论文内容

中文总结 AI 辅助

TDD-Agent将测试驱动开发范式应用于代码生成,通过测试优先推理和迭代双轨优化,在LiveCodeBench和RepoEval上均优于相关基线,提升了代码及测试的有效性。

中文摘要 AI 辅助

大语言模型(LLMs)在代码生成领域已取得显著进展,但在复杂的仓库级任务中仍难以确保正确性。现有方法常将生成的测试作为静态事后验证器,这限制了其对实现过程的指导能力,且当测试本身不完整或不正确时,可能引入误导性反馈。本文提出TDD-Agent,将测试驱动开发范式应用于代码生成:首先提示模型生成可执行测试,鼓励其在实现前明确预期行为;随后利用执行反馈对生成的代码和测试进行迭代双轨优化。我们通过在LiveCodeBench上使用提示变体TDD-prompt,单独验证了测试优先推理的效果,其性能始终优于基于推理的提示基线。基于该发现,我们在仓库级基准RepoEval上评估完整的TDD-Agent框架,结果显示其性能始终优于基于检索的基线和智能体基线。进一步分析表明,迭代优化不仅提升了代码正确性,还增强了生成测试的有效性,使通过率、覆盖率和变异分数均有所提高,这说明测试可作为不断演进的推理 artifacts,而非固定的验证器。我们的源代码可在该https URL获取。

英文摘要

Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging. Existing approaches often use generated tests as static post-hoc validators, which limits their ability to guide implementation and may introduce misleading feedback when the tests themselves are incomplete or incorrect. In this paper, we introduce TDD-Agent, which operationalizes the test-driven development paradigm for code generation. TDD-Agent first prompts the model to generate executable tests, encouraging it to clarify expected behaviors before implementation, and then performs iterative dual-track refinement over both the generated code and tests using execution feedback. We first isolate the effect of test-first reasoning through a prompt variant TDD-prompt on LiveCodeBench, where it consistently improves upon reasoning-based prompting baselines. Building on this finding, we evaluate the full TDD-Agent framework on RepoEval, a repository-level benchmark, and show that it consistently outperforms retrieval-based and agent-based baselines. Additional analyses show that iterative refinement improves not only code correctness but also the effectiveness of the generated tests, yielding higher pass rates, coverage, and mutation scores, suggesting that tests can serve as evolving reasoning artifacts rather than fixed validators. Our source code is available at https://anonymous.4open.science/r/TDD-Agent-Framework-6370/.

发表机构

  • National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
  • School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
  • Qingdao Research Institute and Hangzhou Innovation Institute, Beihang University(北京航空航天大学青岛研究院与杭州创新研究院)
  • School of Software, Beihang University(北京航空航天大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

↑