arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码智能体作为测试套件审计器:在接近官方测试套件检测能力的同时发现官方套件遗漏的问题

Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch

Shuyang Xie, Shuxiao Xie, Feng Zhu, Yanli Ji, Wangmeng Zuo

arXiv 2608.01715首次发表:更新:

发表机构

Harbin Institute of Technology; Fudan University; Sun Yat-sen University(哈尔滨工业大学; 复旦大学; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出编码智能体作为测试套件审计器,结合认证链识别官方套件遗漏的缺陷提交,在接近官方套件覆盖度的同时,无官方套件时其构建的套件表现最优。

AI 中文摘要

在线评测系统(Online Judge)的判定结果,以及基于这些结果构建的数据集和基准,被视为评估和训练代码类大语言模型的 ground truth(基准真值)。然而此前的审计工作已发出警示:官方测试套件会接受存在缺陷的提交代码。不过这些审计仅停留在警示层面,未提供实用的解决方案。本文提出的解决方案包含两部分:一是一个现成的编码智能体,作为测试套件审计器,既构建对抗性测试套件以暴露官方套件遗漏的问题,又在无官方套件的地方提供此类套件;二是一个认证链,用于判定智能体标记的提交是否确实存在缺陷,且不依赖官方评测系统——具体而言,多个独立编写的已通过提交需在每个测试用例上达成预期输出的一致意见,暴力求解方案用于解决分歧,每个问题的验证器则认证每个失败输入的合法性。其中一个此类智能体在 AtCoder 的 20375 份已审计的已通过提交中,识别出 589 份经证实的“已通过但存在缺陷”的提交;将相同认证应用于全部五个智能体后,此类提交的下限总和达 906 份。五个智能体分别评分,在官方套件检测的逻辑缺陷覆盖度上,每个均保持在与官方套件相差 1.7 个百分点以内。在无官方套件可用的截止后 Codeforces 问题上,相同的测试构建方法使五个复现的基准在每个测试输入预算下均领先。在有官方套件的场景中,智能体审计套件的充分性而非默认其有效;在无官方套件的场景中,智能体构建的套件在我们复现并测试的方法中捕获最多的有缺陷提交。

英文摘要

Online-judge verdicts and the datasets and benchmarks built on them are treated as ground truth for evaluating and training large language models for code. Yet prior audits have sounded a warning: official suites accept buggy submissions. These audits, however, stop at the warning and offer no practical remedy. Our remedy has two parts: an off-the-shelf coding agent, serving as a test-suite auditor, both builds adversarial test suites to expose what official suites miss and supplies these suites where no official suite exists; a certification chain determines whether each agent-flagged submission is genuinely buggy without relying on the official judge: multiple independently written accepted solutions agree on the expected output for every test, brute-force solutions settle disagreements, and a per-problem validator certifies each failing input legal. One such agent identifies 589 verified accepted-but-buggy submissions among AtCoder's 20,375 audited accepted submissions; extending the same certification to all five agents yields a union floor of 906 such submissions. Five agents, scored separately, each stay within 1.7pp of official-suite coverage on logic bugs those suites catch. On post-cutoff Codeforces problems with no available official suites, the same test-building method leads all five reproduced baselines at every tested input budget. Where an official suite exists, the agent audits suite adequacy instead of assuming it; where none exists, agent suites catch the most buggy submissions among methods we reproduced and tested.

Comments24 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑