arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

XREPOTEST:面向大语言模型的多语言仓库级单元测试生成基准测试

XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models

Dung Le Quang, Dong Cao Van, Nam Le Hai, Linh Ngo Van, Anh M. T. Bui, Phuong T. Nguyen

arXiv 2608.25939首次发表:更新:

发表机构

Hanoi University of Science and Technology; University of L’Aquila(河内大学科学技术学院; 拉奎拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出XREPOTEST多语言仓库级单元测试生成基准,针对5种少用语言,结合多上下文策略,用14种LLM实验发现独立与仓库级性能差距,为推进现实场景单元测试生成提供支撑。

AI 中文摘要

大语言模型(LLMs)在自动化单元测试生成方面展现出潜力,但现有评估大多依赖独立场景和有限的编程语言,高估了其在现实场景中的可用性。本文提出XREPOTEST,这是一个针对单元测试生成的多语言仓库级基准测试,涵盖Rust、Go、Julia、PHP和Ruby五种未被充分探索的语言。XREPOTEST采用容器化执行框架评估现实仓库约束下的测试,结合文件级、基于LSP的以及检索式等多种上下文增强策略。除测试通过率和覆盖率等标准指标外,本文还提出调用率(IR)以评估生成的测试是否有效验证了预期功能。对包括Claude 4.5、GPT-5.2、DeepSeek V4-Pro和Qwen系列在内的14种最先进LLM的实验显示,独立场景与仓库级性能间存在显著差距,且更丰富的上下文与测试可靠性之间存在权衡。总体而言,XREPOTEST提供了一个具有挑战性且信息丰富的基准测试,以推进现实软件环境中可扩展且鲁棒的单元测试生成,相关数据集和代码可在该公开链接获取。

英文摘要

Large language models (LLMs) have shown promise for automated unit test generation, but existing evaluations largely rely on standalone settings and a narrow set of programming languages, overestimating real-world readiness. We introduce XREPOTEST, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby. XREPOTEST evaluates tests under realistic repository constraints using a containerized execution framework and multiple context augmentation strategies, including file-level, LSP-based, and retrieval-based context. Beyond standard metrics such as test pass rate and coverage, we propose Invocation Rate (IR) to assess whether generated tests meaningfully exercise the intended functionality. Experiments with 14 state-of-the-art LLMs, including Claude 4.5, GPT-5.2, DeepSeek V4-Pro, and Qwen families, reveal a substantial gap between standalone and repository-level performance, as well as trade-offs between richer context and test reliability. Overall, XREPOTEST provides a challenging and informative benchmark to advance scalable and robust unit test generation in realistic software environments. The dataset and code are publicly available at: https://github.com/solis-team/XRepoTest

CommentsAccepted to EMNLP Main 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑