arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文很重要:提高基于大语言模型的单元测试生成的实际可靠性

Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation

Junjie Chen, Ziqi Wang, Lin Yang, Chen Yang, Xiao Chu, Jianyi Zhou, Guangtai Liang, Qianxiang Wang, Dong Wang

arXiv 2607.19682首次发表:更新:

AI 中文总结

研究基于大语言模型的单元测试生成在实际应用中的问题,提出上下文感知工作流程CATGen,通过明确依赖、稳定框架、静态分析等改进设计,经评估在实际项目和基准测试中提升编译成功率与覆盖率,减少生成时间和消耗。

AI 中文摘要

自动化单元测试生成最近受益于大语言模型(LLMs)的进展,但工业部署显示,有前景的研究成果与实际可用性之间仍存在差距。在具有复杂框架和跨文件依赖关系的实际项目中,基于LLM生成的测试经常无法编译,需要昂贵的人工修复,或提供不稳定的覆盖率提升。本文报告了我们在设计、部署和评估CATGen方面的经验,这是一种基于LLM的单元测试生成的上下文感知工作流程。我们发现编译稳健性关键取决于明确项目级依赖关系、稳定测试类框架,并用轻量级静态分析取代基于LLM的迭代修复。这些见解塑造了CATGen的多阶段设计。我们在实际复杂项目的焦点方法和Defects4J基准上评估CATGen,结果表明它显著提高了编译成功率和结构覆盖率,同时减少了生成时间和令牌消耗。结果表明,基于LLM的可靠单元测试生成在实践中更少依赖提示工程,更多依赖基于实际开发约束的系统工程支持。

英文摘要

Automated unit test generation has recently benefited from advances in large language models (LLMs), yet our industrial deployments reveal a persistent gap between promising research results and practical usability. In real-world projects with complex frameworks and cross-file dependencies, LLM-generated tests frequently fail to compile, require costly manual repair, or provide unstable coverage improvements. This paper reports our experience in designing, deploying, and evaluating CATGen, a context-aware workflow for LLM-based unit test generation, informed by repeated industrial failures and refinements. Rather than relying on LLMs to infer incomplete project context, we found that compilation robustness critically depends on making project-level dependencies explicit, stabilizing test class scaffolding, and replacing iterative LLM-based repair with lightweight static analysis. These experience-driven insights shaped CATGen's multi-stage design, which combines structured context retrieval, deterministic test skeleton construction, and program analysis-based post-processing. We evaluate CATGen on real-world complex focal methods from proprietary industrial projects and additionally on the Defects4J benchmark to assess generalizability. Across both settings, CATGen substantially improves compilation success and structural coverage while significantly reducing generation time and token consumption compared to existing LLM-based approaches. Our results demonstrate that reliable LLM-based unit test generation in practice depends less on prompt engineering alone and more on systematic engineering support grounded in real-world development constraints.

Comments23 pages, 5 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑