发表机构
Faculty of Computer Science, University of Vienna; SBA Research gGmbH; Christian Doppler Laboratory AsTra, University of Vienna(维也纳大学计算机学院; SBA研究所; 维也纳大学AsTra克里斯蒂安·多普勒实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文使用本地开放权重语言模型仅凭提交内容标注缺陷修复提交,在手动验证的多语言语料库上显著优于关键词基线,并发布标注流水线与银标准语料库。
AI 中文摘要
缺陷预测依赖于知道哪些提交修复了缺陷,然而编码这一信息的标签是由各种引入噪声的途径产生的。重复使用的基准数据集带有已知的数据质量问题,问题追踪器链接存在偏差,且底层报告经常被错误分类,而在提交消息中匹配关键词是一种粗糙的启发式方法。本文研究是否仅根据提交内容即可将提交标注为缺陷修复,使用本地运行的开放权重语言模型,从而保持过程的可复现性、在语料库规模上成本低廉、可用于专有代码,并且不依赖任何问题追踪器。针对涵盖Java、Python和JavaScript的手动验证和整理的缺陷修复数据集,我们将关键词基线方法与一组不同规模的开放权重模型进行比较,向每个模型输入提交消息和代码差异。在手动验证的语料库上,关键词基线方法找回的修复不足一半,而开放权重模型找回了绝大多数修复,并在逐个仓库的基础上以统计显著性优于基线,且较大的模型并不始终优于较小的模型。我们进一步表明,没有负例的评估语料库无法支持对此类分类器进行精度感知的比较。我们发布了标注流水线以及由推荐配置生成的带标注的多语言语料库,作为构建当前、特定项目数据集的可复现银标准资源。
英文摘要
Defect prediction depends on knowing which commits fix bugs, yet the labels that encode this are produced by routes that each introduce noise. Reused benchmarks carry documented data-quality problems, issue-tracker links are biased and the underlying reports are frequently mistyped, and matching keywords in commit messages is a coarse heuristic. This paper examines whether commits can be labelled as bug fixes from their content alone, using open-weight language models that run locally and therefore keep the process reproducible, inexpensive at corpus scale, usable on proprietary code, and independent of any issue tracker. Against datasets of manually validated and curated bug fixes spanning Java, Python, and JavaScript, we compare a keyword baseline with a set of open-weight models of varying size, prompting each with the commit message and the code diff. On the manually validated corpus the keyword baseline recovers fewer than half of the fixes, whereas the open-weight models recover the large majority and outperform it repository by repository with statistical significance, and larger models do not consistently outperform smaller ones. We further show that evaluation corpora without negative examples cannot support a precision-aware comparison of such classifiers. We release the labelling pipeline together with a labelled, multi-language corpus produced by the recommended configuration, as a reproducible silver-standard resource for building current, project-specific datasets.
Comments19 pages, 6 figures