arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TestJack:你是否应该信任编码基准中的结果?通过评估器演化进行智能体编码基准审计

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

Shuangjie Yao, Hao Wang, Koushik Sen, Simin Chen, Baishakhi Ray, Dawn Song

arXiv 2610.10619首次发表:更新:

AI 中文总结

本研究提出 TestJack 框架,通过审计编码基准中智能体试例的测试合规性,发现约 34.4%被判定正确的试例违反任务要求,凸显现有编码评估器需更具适应性的局限。

AI 中文摘要

大型语言模型(LLM)智能体正在快速重塑软件工程领域,伴随而来的是新编码基准的爆发式增长。然而,几乎所有现有基准仍依赖着沿用数十年的标准:若解决方案能通过固定的一组单元测试,则判定其正确。这类测试往往不够充分:它们仅检查任务要求的部分内容,因此智能体可能通过“ hack”(投机取巧)通过测试,或在通过所有测试的同时悄悄遗漏所需行为。结果是,更高的基准分数可能部分反映了对评估器的更好适配,而非更好的问题解决能力。现有研究聚焦于静态测试增强:它们在任何试例出现前,一次性强化每个任务的测试,从而忽略了实际试例的失败方式。我们提出 TestJack,这是一个用于在固定测试之外评估补丁的可扩展框架。对于每个试例,TestJack 生成针对补丁可能违反的提示要求的测试,仅保留被真实补丁通过的测试,并重新检查任何试例失败。因此,每个确认的失败都有可复现的测试作为支撑。为降低评估成本,我们还引入了一个轻量变体,该变体对随机抽样的试例进行深度审计,并将生成的测试在同一任务的所有试例中复用。在 6 个前沿模型后端和 5 个基准(如 DeepSWE 和 SWE Marathon)上,我们发现当前被判定为正确的模型试例中约有 34.4%违反了任务要求,这使得整体解决率从 50.6%降至 33.2%。我们的结果揭示了当前编码智能体评估的一个根本局限:随着 LLM 更擅长针对固定评估器进行优化,这些评估器本身必须变得更具适应性。

英文摘要

Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on the same decades-old criterion: a solution is correct if it passes a fixed set of unit tests. Such tests are often insufficient: they check only part of what the task requires, so agents can reward hack them or silently miss required behavior while still passing every test. As a result, higher benchmark scores may partly reflect better adaptation to the evaluator rather than better problem solving. Existing works focus on static test augmentation: they strengthen each task's tests once, before any trial is seen, and thus overlook how real trials actually fail. We introduce TestJack, a scalable framework for evaluating patches beyond fixed tests. For each trial, TestJack generates tests targeting prompt requirements the patch may violate, retains only tests passed by the ground-truth patch, and re-examines any trial failures. Each confirmed failure is thus supported by a replayable test. To reduce evaluation cost, we also introduce a lightweight variant which audits a random sample of trials in depth and reuses the resulting tests across all trials for the same task. Across 6 frontier model backends and 5 benchmarks such as DeepSWE and SWE Marathon, we find that about 34.4% of the model trials currently judged correct violate the task requirements, lowering the overall resolution rate from 50.6% to 33.2%. Our results reveal a fundamental limitation of current coding-agent evaluation: as LLMs become better at optimizing against fixed evaluators, those evaluators themselves must become more adaptive.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑