AI 中文总结
研究一对多问题-提交可追溯性,提出LinkRank框架,采用迭代策略,考虑候选提交集,通过构建新数据集评估,在Known-K和Unknown-K设置下均优于基线,表明以问题为中心的排序和迭代选择更优。
AI 中文摘要
恢复问题与提交之间的可追溯性链接对软件维护、调试、影响分析和项目理解很重要。但现有方法多假设一对一关系,忽略一对多关系会导致可追溯性不完整。本文提出LinkRank框架用于恢复一对多问题-提交链接。它考虑问题的候选提交集,与主要独立判断问题-提交对的现有方法不同。为支持实际评估,构建新数据集。LinkRank采用迭代选择-移除-重新归一化策略。在两种设置下评估,结果显示LinkRank在Known-K设置下平均F1分数为74.54%,Unknown-K设置下分别为68.84%(ABS)和67.02%(REL),优于最强基线。研究表明将一对多问题-提交可追溯性作为以问题为中心的排序和迭代选择问题比独立成对分类更好解决。
英文摘要
Recovering traceability links between issues and commits is important for software maintenance, debugging, impact analysis, and project understanding. However, most existing approaches assume a one-to-one relationship, where each issue is linked to a single commit. In practice, many issues are resolved through multiple commits, and ignoring this one-to-many nature can lead to incomplete traceability. This paper presents LinkRank, a learning-to-rank framework for recovering one-to-many issue--commit links. Unlike existing methods that mainly judge issue--commit pairs independently, LinkRank considers the set of candidate commits for an issue and identifies the commits that are most likely to contribute to its resolution. To support realistic evaluation, we construct a new dataset from six open-source GitHub repositories. LinkRank follows an iterative pick--remove--renormalize strategy: it selects the highest-ranked commit, removes it from the candidate pool, renormalizes the remaining scores, and repeats the process until the stopping criterion is met. We evaluate LinkRank under two settings: Known-K, where the true number of linked commits is available, and Unknown-K, where the model must infer when to stop selecting commits using ABS and REL stopping rules. Across six projects, LinkRank achieves an average Known-K F1 score of 74.54%, compared with 48.39% for the strongest baseline. In the Unknown-K setting, LinkRank achieves 68.84% F1 with ABS and 67.02% F1 with REL, outperforming the strongest baselines under both automatic stopping rules. Overall, the findings suggest that one-to-many issue--commit traceability is better addressed as an issue-centric ranking and iterative selection problem than as independent pairwise classification.