arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TestPrism:重新审视超越单一参考的测试评估

TestPrism: Rethinking Test Evaluation Beyond a Single Reference

Han Li, Lingxiang Hu, Jiacheng Huang, Ziqian Jiang, Jingkai Luo, Wei Gao, Yunfan Tan, Zun Wang, Jiaheng Liu

arXiv 2610.12289首次发表:更新:

发表机构

Nanjing University; Hong Kong Polytechnic University; Tencent(南京大学; 香港理工大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TestPrism针对LLM编码智能体测试评估的单一参考缺陷构建,提出联合成功函数指标,推出TestHelix方法提升测试评估性能。

AI 中文摘要

大型语言模型(LLM)编码智能体已在各类编程任务的测试生成方面取得进展。然而,针对单一参考解决方案评估测试的常规做法,忽略了其他有效的实现方式,可能会高估测试质量。我们推出TestPrism,包含来自17个来源的300个测试任务和3000个候选实现,其中有效和无效解决方案各占一半。其核心指标联合成功函数要求生成的测试在初始程序状态下失败,接受所有有效候选,拒绝所有无效候选。在14种基线编码智能体配置中,联合成功函数仅达到28.00%,而单一参考成功达到59.67%。我们的分析发现了缺失行为、不支持的断言和有缺陷的测试构建。为解决这些弱点,我们推出TestHelix,它将测试与修复对的异构合成、同行交叉验证以及递归自改进(RSI)相结合。在两个模型上,TestHelix在TestHelix评估中比原生工具比较器将联合成功函数提高了8.67至9.00个百分点。

英文摘要

Large language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference solution overlooks alternative valid implementations and can overstate test quality. We introduce TestPrism, comprising 300 test tasks from 17 sources and 3000 candidate implementations, evenly split between valid and invalid solutions. Its primary metric, Joint Success Function, requires the generated tests to fail on the initial program state, accept every valid candidate, and reject every invalid candidate. Across fourteen baseline coding agent configurations, Joint Success Function reaches only 28.00%, whereas single reference success reaches 59.67%. Our analysis reveals missed behaviors, unsupported assertions, and faulty test construction. To address these weaknesses, we introduce TestHelix, which combines heterogeneous synthesis of test and repair pairs with peer cross validation and recursive self improvement (RSI). Across two models, TestHelix improves Joint Success Function by 8.67 to 9.00 percentage points over the native harness comparators in the TestHelix evaluation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑