arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17993cs.SE

基于LLM的程序修复中的实验设置:关于输入、工具访问、反馈和验证的研究

Experimental Settings in LLM-Based Program Repair: A Study of Inputs, Tool Access, Feedback, and Validation

Xushu Dai, Yicheng Cai, Nanqing Luo, Pei-Yu Tseng

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对基于LLM的程序修复实验设置不明确的问题,提出显式规范框架和机器可读模式,以提升可复现性和跨系统比较的清晰度。

中文摘要 AI 辅助

自动程序修复(APR)系统的评估通常报告基准、修复缺陷的数量以及用于最终补丁验证的测试,但这些项目不再完全指定呈现给系统的修复任务。最近的基于LLM的系统在修复前提供的信息、修复期间允许的仓库和测试操作以及失败尝试后返回的反馈方面存在差异,使得同一基准可以实例化从局部补丁生成到仓库级诊断和迭代修复等实质不同的修复任务。我们提出了一个框架,用于明确指定与所报告的APR结果相关的实验设置。我们分析了在Defects4J和SWE-bench上评估的系统的报告实验设置,并通过任务单元、故障定位假设、初始输入、工具访问、修复时反馈、最终验证和资源预算来表征每个结果。我们的分析表明,仅凭基准身份不足以重建被评估的任务或确定所报告修复率之间比较的适当范围。因此,我们引入了一种机器可读的模式来指定每个实验设置,以提高可复现性并使跨系统比较的范围明确化。

英文摘要

Evaluations of automated program repair (APR) systems commonly report the benchmark, the number of repaired defects, and the tests used for final patch validation, but these items no longer fully specify the repair task presented to a system. Recent LLM-based systems differ in the information supplied before repair, the repository and testing operations permitted during repair, and the feedback returned after unsuccessful attempts, allowing the same benchmark to instantiate substantially different repair tasks ranging from localized patch generation to repository-level diagnosis and iterative repair. We present a framework for explicitly specifying the experimental settings associated with reported APR results. We analyze reported experimental settings from systems evaluated on Defects4J and SWE-bench and characterize each result by its task unit, fault-localization assumptions, initial input, tool access, repair-time feedback, final validation, and resource budget. Our analysis shows that benchmark identity alone is insufficient to reconstruct the evaluated task or determine the appropriate scope of comparison across reported repair rates. We therefore introduce a machine-readable schema for specifying each experimental setting to improve reproducibility and make the scope of cross-system comparisons explicit.

发表机构

  • Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑