arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28896cs.SE

规范驱动的自动程序修复基准测试:从静态语料库到可执行规范

Specification-Driven Benchmarking for Automated Program Repair From Static Corpora to Executable Specifications

Yasser Ebrahim

首次发表
浏览论文内容

中文总结 AI 辅助

提出规范驱动的基准测试范式,用可执行规范定义基准并通过生成流水线实现,实现基准的声明式实验设计。

中文摘要 AI 辅助

自动程序修复(APR)基准传统上被构建为静态数据集,其特性继承自所包含的缺陷。虽然这一范式推动了数十年的进展,但有限的语料库提供的实验控制有限,随着重复使用而越来越容易受到污染,并且无法随着评估需求的发展而系统性地重新生成或调整。我们提出规范驱动的基准测试,这是一种通过可执行规范定义基准并通过基准生成实现的范式。规范明确声明了基准的预期属性(包括程序上下文、故障分类、难度、验证策略和语料库约束),而生成流水线通过独立的生成、验证和语料库管理组件实现这些要求。我们通过引入基准规范维度的分类法来发展这种方法的概念基础,建立每个规范维度如何映射到确定性的架构职责,并论证独立验证是可信基准生成的结构性要求。一个端到端的示例说明了规范选择如何通过流水线传播,以生成其属性可独立验证的基准实例。通过将基准视为可执行规范而非静态数据集,所提出的范式将基准构建从工件管理转变为声明式实验设计。

英文摘要

Automated Program Repair (APR) benchmarks have traditionally been constructed as static datasets whose characteristics are inherited from the defects they contain. While this paradigm has enabled decades of progress, finite corpora provide limited experimental control, become increasingly susceptible to contamination as they are reused, and cannot be systematically regenerated or adapted as evaluation requirements evolve. We propose specification-driven benchmarking, a paradigm in which benchmarks are defined by executable specifications and realized through benchmark generation. The specification explicitly declares the intended properties of the benchmark (including program context, fault taxonomy, difficulty, validation strategy, and corpus constraints) while a generation pipeline realizes those requirements through independent generation, validation, and corpus management components. We develop the conceptual foundations of this approach by introducing a taxonomy of benchmark specification dimensions, establishing how each specification dimension maps to deterministic architectural responsibilities, and arguing that independent validation is a structural requirement for trustworthy benchmark generation. An end-to-end example illustrates how specification choices propagate through the pipeline to produce benchmark instances whose properties are independently verifiable. By treating the benchmark as an executable specification rather than a static dataset, the proposed paradigm shifts benchmark construction from artifact curation to declarative experimental design.

发表机构

  • Algoma University(阿尔戈马大学)

机构由 AI 辅助整理,请以论文原文为准。

↑