arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Multi-SWT-Bench:用于复现测试生成的多语言基准

Multi-SWT-Bench: A Multilingual Benchmark for Reproduction Test Generation

Kazuki Kusama, Sota Nakashima, Haruka Tokumasu, Masanari Kondo, Lingming Zhang, Yasutaka Kamei

arXiv 2609.34752首次发表:更新:

发表机构

Kyushu University; University of Illinois Urbana-Champaign(九州大学; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有复现测试生成基准仅限单语言的问题,提出多语言基准Multi-SWT-Bench,含8种语言1963个实例,实证发现语言差距,为跨生态系统的测试生成方法提供方向。

AI 中文摘要

复现测试生成将自然语言的问题描述转化为可执行的测试,这些测试在原始代码上失败,在问题解决后通过,从而为验证候选补丁提供可执行的证据。现有基准是针对个别编程语言构建的,阻碍了跨多样化编程生态系统的统一评估。为解决这一局限,我们引入了MULTI-SWT-BENCH,一个用于复现测试生成的多语言基准,包含跨八种编程语言(Python、Java、TypeScript、JavaScript、Go、Rust、C和C++)的1,963个实例。利用该基准,我们使用四种代表性方法(MSWE-agent、MOpenHands、Codex和Claude Code)对最先进的大语言模型进行了实证研究,并跨编程语言进行了失败分析。我们的评估揭示了系统性的语言差距。在每种评估方法和LLM中,Python上的成功率均超过所有语言的总体成功率,而C++的成功率尤其低。我们的失败分析识别了源于仓库测试约定的语言特定挑战,以及推断隐式设置要求和通过迭代修订保持目标行为方面的跨语言挑战。这些发现证明了多语言评估的重要性,并为开发能跨软件生态系统泛化并可靠捕获问题特定行为的复现测试生成方法提供了可行方向。

英文摘要

Reproduction test generation translates a natural-language issue description into executable tests that fail on the original code and pass after the issue is resolved, providing executable evidence for verifying candidate patches. Existing benchmarks are constructed for individual programming languages, preventing a unified evaluation across diverse programming ecosystems. To address this limitation, we introduce MULTI-SWT-BENCH, a multilingual benchmark for reproduction test generation consisting of 1,963 instances across eight programming languages: Python, Java, TypeScript, JavaScript, Go, Rust, C, and C++. Using this benchmark, we conduct an empirical study of state-of-the-art LLMs with four representative methods (MSWE-agent, MOpenHands, Codex, and Claude Code) and perform a failure analysis across programming languages. Our evaluation reveals a systematic language gap. Across every evaluated method and LLM, the success rate on Python exceeds the aggregate success rate across all languages, while C++ exhibits particularly low success rates. Our failure analysis identifies both language-specific challenges arising from repository testing conventions and cross-language challenges in inferring implicit setup requirements and preserving the target behavior through iterative revisions. These findings demonstrate the importance of multilingual evaluation and provide actionable directions for developing reproduction test generation methods that generalize across software ecosystems and reliably capture issue-specific behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑