arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型的测试生成中的代码健康度:有效性与标记效率

Code Health in LLM-Based Test Generation: Effectiveness and Token Efficiency

Freya Wirdemann, Markus Borg, Nadim Hagatulah, Adam Tornhill

arXiv 2608.18645首次发表:更新:

AI 中文总结

本研究探究LLM生成单元测试的有效性随CodeScene的CodeHealth水平的变化,发现CH是测试有效性的微弱但一致信号,且与输入标记数负相关,证实可维护性与LLM软件开发存在关联。

AI 中文摘要

由大语言模型(LLM)驱动的编码智能体如今在软件工程领域已十分突出。过往研究表明,AI工具在易于维护的高质量源代码上表现更佳。本研究探究了LLM生成的单元测试的有效性如何随CodeScene的CodeHealth(CH)所衡量的可维护性水平变化而变化;我们采用传统的覆盖率指标和变异分数,在Python、Java和C++三种语言上评估测试有效性,还研究了不同CH水平的代码在使用常见工业级标记器时转换为输入标记的情况。结果显示,CH对LLM生成测试的有效性提供了微弱但一致的信号,且与输入标记数呈负相关;这些发现进一步证明了可维护性与基于LLM的软件开发之间存在关联。

英文摘要

Coding agents powered by Large Language Models (LLMs) are now prominent in software engineering. Previous work has shown that AI tools perform better on high-quality source code that is easy to maintain. In this study, we investigate how the effectiveness of LLM-generated unit tests varies across maintainability levels measured by CodeScene's CodeHealth (CH). We assess test effectiveness using traditional coverage metrics and mutation score across Python, Java, and C++. Moreover, we study how code with different levels of CH translates into input tokens using common industrial tokenizers. Our results suggest that CH provides a weak but consistent signal of LLM-generated test effectiveness and is negatively correlated with input-token count. These findings provide further evidence for a relationship between maintainability and LLM-based software development.

CommentsAccepted at the Engineering Track of the 26th IEEE International Conference on Source Code Analysis and Manipulation (SCAM 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑