AI 中文总结
研究针对影响分析这一软件维护任务,提出Athena方法,结合软件系统依赖图信息与概念耦合方法,无需更改历史和执行信息。构建新基准测试,实验表明该方法在新基准上表现出色,比简单基线有显著性能提升。
AI 中文摘要
影响分析(IA)是一项关键的软件维护任务,旨在识别给定代码更改对大型软件项目的影响,避免潜在负面影响。它具有认知挑战性,涉及推理各种代码结构之间的抽象关系。此前许多基于耦合度量的IA方法存在不足。本文介绍了一种名为Athena的新颖IA方法,它将软件系统的依赖图信息与概念耦合方法相结合,无需更改历史和执行信息。之前的IA基准较小且存在问题,因此构建了名为Alexandria的大规模基准。在新基准上,最佳方法配置的mRR、mAP和HIT@10分数分别为60.32%、35.19%和81.48%。通过各种分析表明,Athena的程序依赖图和概念耦合信息的新颖组合使其比简单基线有显著性能提升。
英文摘要
Impact analysis (IA) is a critical software maintenance task that identifies the effects of a given set of code changes on a larger software project with the intention of avoiding potential adverse effects. IA is a cognitively challenging task that involves reasoning about the abstract relationships between various code constructs. Given its difficulty, researchers have worked to automate IA with approaches that primarily use coupling metrics as a measure of the "connectedness" of different parts of a software project. Many of these coupling metrics rely on static, dynamic, or evolutionary information and are based on heuristics that tend to be brittle, require expensive execution analysis, or large histories of co-changes to accurately estimate impact sets. In this paper, we introduce a novel IA approach, called Athena, that combines a software system's dependence graph information with a conceptual coupling approach that uses advances in deep representation learning for code without the need for change histories and execution information. Previous IA benchmarks are small, containing fewer than ten software projects, and suffer from tangled commits, making it difficult to measure accurate results. Therefore, we constructed a large-scale IA benchmark, called Alexandria, from 25 open-source software projects, that utilizes fine-grained commit information from bug fixes. On this new benchmark, our best-performing approach configuration achieves mRR, mAP, and HIT@10 scores of 60.32%, 35.19%, and 81.48%, respectively. Through various ablations and qualitative analyses, we show that Athena's novel combination of program dependence graphs and conceptual coupling information leads it to outperform a simpler baseline by 10.34%, 9.55%, and 11.68% with statistical significance.
CommentsAccepted to FSE'24
DOI:10.1145/3643770