发表机构
University of California, San Diego; Purdue University; National University of Singapore(加州大学圣地亚哥分校; 普渡大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WitnessGym通过错误注入自动构建错误验证基准,生成1,300个Java案例,评估显示编码智能体在见证构建任务上仍面临挑战。
AI 中文摘要
错误验证要求编码智能体为所报告的错误生成一个可执行的见证。该见证结合了具体输入与测试框架,并在执行过程中暴露错误行为。此类证据使审计发现具有可操作性,然而当案例复用公开的历史错误和见证或需要人工构建时,基准评估变得困难。我们提出了WitnessGym,一个通过错误注入构建错误验证基准的自动化框架。它将错误注入到真实项目的测试可达路径中,重建每个项目,并保留由构建时见证暴露的案例。错误规范和执行适配器允许扩展到其他错误类型和语言。错误保持变换在保持见证行为的同时改变周围结构。基于带有测试套件的真实世界Java项目,WitnessGym自动构建了1,300个基准案例。注入的补丁与历史错误补丁相似,且两个被评估的模型在盲比较中难以区分。我们在六种框架/模型配对中评估了四个编码智能体框架,涵盖错误类型、执行上下文和变换深度。即使已知错误模式,见证构建仍然困难。我们的框架、基准案例和评估脚本均已公开。
英文摘要
Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.
Comments30 pages, 8 figures, and 12 tables, including appendices