发表机构
Bilkent University(比尔肯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过微调GPT-4o及多款开源LLM解决PR-问题对齐分类问题,利用SHAP分析关键影响因素,实现准确率等指标提升,CodeLlama-7B表现最佳。
AI 中文摘要
背景:拉取请求(PR)与对应问题的准确对齐对高效软件开发和维护代码质量至关重要,对齐错误会降低可追溯性、阻碍缺陷定位并降低可维护性。目的:本研究旨在利用微调后的大型语言模型(LLM)改进跨多种对齐类别的自动化PR-问题对齐分类,并开展可解释性分析以探究PR-问题字段对微调后LLM预测的影响。方法:我们的方法包括数据集准备、LLM微调及可解释性分析。首先,我们扩展了现有数据集并应用数据增强以解决类别不平衡问题;随后通过指令微调对GPT-4o进行微调,对包括CodeLlama-7B、CodeQwen1.5-7B、StableCode-3B、CodeGemma-7B和Deepseek-Coder-6.7B在内的开源LLM使用分类专用头进行微调;采用Shapley加性解释(SHAP)开展可解释性分析,以探究PR-问题字段对性能最佳的开源LLM预测的影响。结果:微调后的LLM优于基线模型,在准确率和F1-micro上平均提升6.15%,F1-macro提升14.69%,召回率提升6.15%;CodeLlama-7B成为整体性能最佳的微调LLM,可解释性分析显示代码差异、问题主体及PR主体内容对预测的影响最大。结论:微调可显著增强PR-问题对齐分类,提升准确率和效率;可解释性分析为驱动对齐决策的数据集特征提供了可操作的见解,加深了对LLM如何推理软件制品的理解。
英文摘要
Context: Accurate alignment between pull requests (PRs) and corresponding issues is crucial for efficient software development and maintaining code quality, as misalignments can reduce traceability, hinder defect localization, and decrease maintainability. Objective: This study aims to improve automated PR-issue alignment classification by leveraging fine-tuned large language models (LLMs) across multiple alignment categories, and conducts interpretability analysis to investigate the effects of PR-issue fields on the predictions of fine-tuned LLMs. Method: Our methodology consists of dataset preparation, LLM fine-tuning, and interpretability analysis. We first extended an existing dataset and applied data augmentation to address class imbalance. GPT-4o was then fine-tuned via instruction tuning, and open-source LLMs including CodeLlama-7B, CodeQwen1.5-7B, StableCode-3B, CodeGemma-7B, and Deepseek-Coder-6.7B were fine-tuned using classification-specific heads. Interpretability analysis using Shapley Additive Explanations (SHAP) was conducted to examine the influence of PR-issue fields on predictions for the best-performing open-source LLM. Results: Fine-tuned LLMs outperformed baseline models, achieving average improvements of 6.15% in accuracy and F1-micro, 14.69% in F1-macro, and 6.15% in recall. CodeLlama-7B emerged as the best-performing fine-tuned LLM overall, while interpretability analysis revealed that code diffs together with issue body and PR body contents exert the greatest influence on predictions. Conclusions: Fine-tuning substantially enhances PR-issue alignment classification, improving both accuracy and efficiency. Interpretability analysis provides actionable insights into the dataset features driving alignment decisions, deepening understanding of how LLMs reason over software artifacts.
Comments31 pages, 12 figures. Submitted to Springer Empirical Software Engineering (EMSE), Special Issue on Software Analysis, Evolution, and Reengineering (SANER 2025)