arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结合问题信息的大语言模型增强型提交消息生成:一项探索性研究

LLM-Enhanced Commit Message Generation via Issue Information: An Exploratory Study

Zongen Ren, Wei Shi, Bo Xiong, Chong Wang, Peng Liang

arXiv 2608.22004首次发表:更新:

AI 中文总结

该研究提出ISAC框架,将代码差异与问题信息结合作为LLM输入生成提交消息,经实验验证其性能优于现有基线,且纳入问题信息可提升CMG效果。

AI 中文摘要

提交消息可帮助开发者理解代码变更、支持协作并提升长期维护效率。然而,仅将问题信息作为基于大语言模型(LLM)的提交消息生成(CMG)的外部上下文,尚未得到系统研究。我们提出一种问题增强型提交消息生成框架ISAC(ISsue-Augmented framework for Commit message generation),将代码差异与问题信息结合作为LLM的输入。为支撑评估,我们构建了ApacheCM-Issue数据集,该数据集基于ApacheCM构建,通过将提交与GitHub及Apache Jira中的问题关联,实现提交-问题对齐。我们使用Scala、Java和C++项目的样本,在不同推理配置下,采用两款代表性LLM(GPT-5.5和DeepSeek-V4-Flash)评估四种输入配置。结果显示,在所有评估的模型配置和指标中,纳入问题信息均能持续提升基于LLM的CMG性能,其中CIDEr指标的提升最为显著。纳入类似历史提交可进一步提升自动指标得分,而用结构化问题摘要替代完整问题信息则会降低得分。在实验数据集上,ISAC在全部五项自动指标上均优于四个复现的最先进(SOTA)CMG基线。人工评估进一步表明,结构化问题摘要可提升感知完整性,但替换原始问题信息会牺牲上下文细节,导致自动指标结果变差。

英文摘要

Commit messages help developers understand code changes, support collaboration, and improve long-term maintenance. However, the use of issue information alone as the external context for LLM-based CMG has not been systematically studied. We propose an ISsue-Augmented framework for Commit message generation (ISAC) by combining code diffs with issue information as LLM input. To support the evaluation, we construct ApacheCM-Issue, a commit-issue aligned dataset built upon ApacheCM by linking commits with issues from GitHub and Apache Jira. Using samples from Scala, Java, and C++ projects, we evaluate four input configurations using two representative LLMs, GPT-5.5 and DeepSeek-V4-Flash in different reasoning configurations. The results show that incorporating issue information consistently improves LLM-based CMG across all evaluated model configurations and metrics, with the largest gains observed for CIDEr. Incorporating a similar historical commit further improves automatic metric scores, while replacing full issue information with a structured issue summary decreases them. ISAC also outperforms the four reproduced state-of-the-art (SOTA) CMG baselines across all five automatic metrics on the experimental dataset. The human evaluation further shows that structured issue summaries may improve perceived completeness, although replacing the original issue information can sacrifice contextual details and lead to worse results on automatic metrics.

Comments28 pages, 7 images, 11 tables, Manuscript submitted to a journal (2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑