arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

故障定位是否优于全新尝试?一项测试引导代码修复的安慰剂对照研究

Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair

Anik Jha

arXiv 2609.00854首次发表:更新:

AI 中文总结

该研究通过安慰剂对照实验,对比盲重采样、定位填充等方法,发现测试引导代码修复中故障定位效果逊于盲采样,仅总体略优于随机区间填充,且仅适用于24-32B模型。

AI 中文摘要

故障定位可将代码模型的修复聚焦于失败测试所关联的语句,但针对性编辑可能仅因规模较小而成功,且第二次模型调用可能完全不利用失败信息便成功。我们对同一失败候选方案采用三种处理方式以区分上述解释:盲全方案重采样、基于谱的定位后可疑区间填充、在不相交随机代码区间进行等长填充。实验使用三个冻结的26-32B模型、三个基准及488个失败候选方案,外加第三个系列中单独声明的24B第四模型,得到三个结果。第一,定位极少可用:仅9.0%的失败候选方案暴露带有可用谱的失败公开测试。第二,在强测试套件可定位的177个候选方案中,匹配尝试次数下,定位填充明显逊于盲重采样(3:40,p=3.0×10^-9),与我们的假设相反;该差异在第三个系列中以-11.3个百分点(95%置信区间[-16.6,-6.8])复现,且扩大编辑规模无法挽回劣势。第三,与随机区间安慰剂相比,定位填充总体占优(11:1,霍尔姆校正后p=0.019),但在我们部署计划指定的主要分析下,该优势在单个模型中均不显著(最佳霍尔姆p=0.087),因此我们将定位效应报告为提示性而非确定。将尝试重新定价为token会缩小但不推翻此结果:区间尝试花费21.7个生成token,而盲尝试花费371.1个,然而16次定位尝试达到6.8%,而1次盲尝试已达到10.1%。48.9%的尝试会逐字重现被移除的区间,这便是更多预算无济于事的原因。我们将所有定位结论限制于所测试的24-32B模型。

英文摘要

Fault localization can focus a code model's repair on the statements a failing test implicates, but a targeted edit may succeed merely because it is small, and a second model call may succeed without using the failure at all. We separate these explanations with three arms applied to the same failed candidate: blind whole-solution resampling, spectrum-based localization followed by suspect-span infilling, and same-length infilling at a disjoint random code span. Across three frozen 26-32B models, three benchmarks and 488 failing candidates, plus a separately declared 24B fourth model from a third family, three results follow. First, localization is rarely available: only 9.0% of failing candidates expose a failing public test with a usable spectrum. Second, among the 177 candidates localizable from a strong suite, localized infilling loses decisively to blind resampling at a matched attempt count (3:40, p = 3.0 x 10^-9), opposite to our hypothesis; the loss replicates in a third family at -11.3 points (95% CI [-16.6, -6.8]), and widening the edit does not rescue it. Third, against the random-span placebo localized infilling leads pooled (11:1, Holm-adjusted p = .019), but that lead resolves in no individual model under the analysis our shipped plan designates primary (best Holm p = .087), so we report the location effect as suggestive rather than established. Re-pricing attempts as tokens narrows but does not overturn this: a span attempt spends 21.7 generated tokens against 371.1, yet 16 localized attempts reach 6.8% while one blind attempt already reaches 10.1%. Infilling reproduces the removed span verbatim in 48.9% of attempts, which is why more budget does not help. We restrict every localization conclusion to the 24-32B models tested.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑