发表机构
University of Illinois at Chicago(伊利诺伊大学香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对开发人员编写的信息不足的Git提交消息问题,提出CommitLLM管道,通过微调小语言模型、约束解码和后处理,从代码差异生成合规消息,经评估效果良好,后处理对质量提升贡献更大。
AI 中文摘要
开发人员经常编写诸如“修复”或“更新内容”等信息不足的Git提交消息,这降低了代码审查、调试和入职培训中版本控制历史的价值。我们提出了CommitLLM,这是一个三阶段管道,它使用微调后的小语言模型从代码差异中生成简洁的、符合常规提交规范的消息。该系统结合了(1)在CommitPackFT数据集上对Mistral-7B-Instruct-v0.2进行QLoRA微调,(2)使用约束解码来确保简洁性,以及(3)确定性后处理以去除对话式工件并强制格式。在50个样本的评估中,CommitLLM实现了98%的格式合规性(普通Mistral为22%),将平均输出长度从154.8个字符减少到37.9个字符,并将LLM作为评判的分数从1.97提高到3.68(满分5分)。值得注意的是,后处理层对质量提升的贡献比微调本身更大,这表明对于结构化输出任务,将语言模型视为确定性管道中的一个组件比单独优化模型更有效。整个系统在单个消费级GPU(NVIDIA T4,16GB VRAM)上运行。
英文摘要
Developers frequently write uninformative git commit messages such as "fix" or "update stuff", degrading the value of version-control history for code review, debugging, and onboarding. We present CommitLLM, a three-stage pipeline that generates concise, Conventional Commits-compliant messages from code diffs using a fine-tuned small language model. The system combines (1) QLoRA fine-tuning of Mistral-7B-Instruct-v0.2 on the CommitPackFT dataset, (2) constrained decoding to enforce brevity, and (3) deterministic post-processing to strip conversational artifacts and enforce format. On a 50-sample evaluation, CommitLLM achieves 98% format compliance (vs. 22% for vanilla Mistral), reduces average output length from 154.8 to 37.9 characters, and improves LLM-as-a-Judge scores from 1.97 to 3.68 out of 5. Notably, the post-processing layers contribute more to quality improvement than the fine-tuning itself, suggesting that for structured-output tasks, treating the LLM as a component in a deterministic pipeline is more effective than optimizing the model alone. The entire system runs on a single consumer GPU (NVIDIA T4, 16 GB VRAM).
Comments7 pages, 4 figures