发表机构
Queen’s University; Université du Québec à Montréal(女王大学; 蒙特利尔魁北克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LabelMate是基于LLM的框架,可从历史问题报告生成项目专属标签集,无需预标注数据即可自动标注新报告,在30个GitHub仓库的16500份报告上准确率达89.84%,优于现有通用方法。
AI 中文摘要
软件用户常向产品的问题跟踪系统提交问题报告,以报告缺陷、提出增强功能建议或反馈其他产品相关问题。对这些问题报告进行标注有助于高效规划并提升社区参与度,但由于设计合适的标注分类体系并为新问题报告分配对应标签需大量人工工作,许多问题报告仍处于未标注状态。现有自动标注方法虽试图缓解这些挑战,但存在关键局限,例如需要大量人工干预、分配通用标签、依赖现有标注数据集等。为解决这些局限,本文提出LabelMate,一种新型基于大语言模型(LLM)的框架,该框架可从历史问题报告中推导生成全面的、针对特定项目的标签集,且无需任何预标注训练数据即可自动为新问题报告分配相关标签。我们在来自30个流行且多样化的GitHub仓库的16500份问题报告上对LabelMate进行评估,基于该数据集,我们的方法生成了275个标签的连贯列表,实现了89.84%的平均标注准确率,相比现有通用标签分配方法具有统计学意义的显著提升。这些结果表明,LabelMate提供了一种高效、领域自适应的解决方案,可简化问题标注流程。
英文摘要
Software users often submit issue reports to a product's issue tracking system to report defects, suggest enhancements, or raise other product-related concerns. Labeling these issue reports supports effective planning and improves community engagement. However, many issue reports remain unlabeled due to the substantial manual effort required to design an appropriate label taxonomy, then assign suitable labels from this taxonomy to new issue reports. Existing automated labeling approaches attempt to mitigate these challenges. However, they suffer from key limitations, such as extensive manual intervention, the assignment of generic labels, and a dependence on existing labeled datasets. To address these limitations, we propose LabelMate, a novel Large Language Model (LLM)-driven framework that (1) derives a comprehensive, project-specific label set from historical issue reports and (2) automatically assigns relevant labels to new issue reports without requiring any pre-labeled training data. We evaluate LabelMate on 16,500 issue reports from 30 popular and diverse GitHub repositories. Based on this dataset, our approach generates a coherent list of 275 labels and achieves an average labeling accuracy of 89.84%, a statistically significant improvement over existing generic label assigning approaches. These results demonstrate that LabelMate offers an efficient, domain-adaptive solution to streamline the issue labeling process.
Comments38 pages, 11 figures