VibeCheck:评估LLM生成的单元测试质量:跨异构仓库的多智能体实证研究
VibeCheck: Assessing the Quality of LLM-Generated Unit Tests: A Multi-agent Empirical Study across Heterogeneous Repositories
浏览论文内容
中文总结 AI 辅助
VibeCheck通过多智能体实证研究,评估LLM生成的单元测试质量,发现测试可运行但断言弱、边界覆盖不足,提出可靠性导向的评估框架。
中文摘要 AI 辅助
基于LLM的IDE智能体越来越多地被用于生成基于仓库的单元测试,然而常见的评估往往依赖于执行成功或覆盖率。这些指标可能遗漏更深层的质量问题,如断言薄弱、边界情况缺失、隔离性差和可维护性有限。本文介绍了VibeCheck,一项针对15个学生开发的Python和JavaScript/TypeScript仓库的单元测试生成的实证研究。我们评估了Kiro、Antigravity和Cursor,以Claude Sonnet 4.5作为底层智能体,在仅仓库、零样本条件下,使用涵盖可运行性、断言强度、逻辑和边界情况覆盖、隔离性/确定性以及可维护性的五维评分标准。我们还应用留一法跨智能体同行评估来比较工具并识别失败模式。结果显示存在明显的执行充分性差距:生成的测试通常可运行,但经常缺乏强断言和有意义的行为覆盖。弱断言和缺失边界情况比阻塞性失败更常见,表明可运行的测试仍然可能是浅层的。VibeCheck为超越通过/失败结果评估LLM生成的测试提供了一个面向可靠性的框架。
英文摘要
LLM-based IDE agents are increasingly used to generate repository-grounded unit tests, yet common evaluations often rely on execution success or coverage. These metrics can miss deeper quality issues such as weak assertions, missing edge cases, poor isolation, and limited maintainability. This paper presents VibeCheck, an empirical study of unit test generation across 15 student-developed Python and JavaScript/TypeScript repositories. We evaluate Kiro, Antigravity, and Cursor with Claude Sonnet 4.5 as the underlying agent, under repository-only, zero-shot conditions using a five-dimensional rubric covering runnability, assertion strength, logic and edge-case coverage, isolation/determinism, and maintainability. We also apply leave-one-out cross-agent peer evaluation to compare tools and identify failure patterns. Results show a clear execution-adequacy gap: generated tests are often runnable but frequently lack strong assertions and meaningful behavioral coverage. Weak assertions and missing edge cases occur more often than blocking failures, showing that runnable tests can still be shallow. VibeCheck provides a reliability-oriented framework for evaluating LLM-generated tests beyond pass/fail outcomes.
发表机构
- University of Dhaka(达卡大学)
机构由 AI 辅助整理,请以论文原文为准。