arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

平衡有用性和自然性:基于大语言模型的代码审查评论整理管道

Balancing Usefulness and Naturalness: An LLM-based Curation Pipeline for Code Review Comments

Oussama Ben Sghaier, Martin Weyssow, Houari Sahraoui

arXiv 2607.09524首次发表:更新:

AI 中文总结

研究针对现有代码审查数据集质量问题限制大语言模型在代码审查任务中效用的情况,提出两种整理管道。一种用大语言模型重述评论生成CuREV数据集,另一种以高质量示例为指导增强评论真实性和多样性,提升了代码审查数据集质量和下游任务效果。

AI 中文摘要

代码审查是软件开发的基石,审查者通过书面评论提供反馈以确保代码质量、可维护性和正确性。此过程的有效性取决于评论质量。随着大语言模型在自动化代码审查任务中受到关注,其效用受训练数据集质量直接限制。现有代码审查数据集常嘈杂、不一致或结构不佳,阻碍大语言模型学习生成准确、有用且类人的审查评论。为克服这些限制,我们提出两种不同的整理管道,旨在提高大规模代码审查数据集的质量和效用。在第一个管道中,大语言模型系统地重新表述所有评论以提高清晰度、简洁性和文明性,同时保留语义意图。由此产生的整理数据集CuREV提供更清晰、高质量且易于学习的评论,可在下游自动化任务中带来可衡量的改进。在此基础上,我们提出一种改进管道,以高质量示例为指导,增强整理评论的真实性和多样性。该方法首先根据使用评估框架的系统质量评估将数据集分为高质量和低质量评论。高质量评论保持原始形式并用作上下文示例以启发低质量评论的重新表述。通过改变提供的示例,重新表述的评论不仅更清晰、更具可操作性,而且展现出更广泛的写作风格,使其更真实、更类人。

英文摘要

Code review is a cornerstone of software development, where reviewers provide feedback through written comments to ensure code quality, maintainability, and correctness. The effectiveness of this process hinges on the quality of review comments. As large language models (LLMs) gain traction in automating code review tasks, the utility of these systems is directly limited by the quality of the datasets on which they are trained. Unfortunately, existing code review datasets are often noisy, inconsistent, or poorly structured, which hinders the ability of LLMs to learn to generate accurate, helpful, and human-like review comments. To overcome these limitations, we propose two different curation pipelines designed to improve both the quality and the utility of large-scale code review datasets. In the first pipeline, all review comments are systematically reformulated by an LLM to improve their clarity, conciseness, and civility while preserving their semantic intent. The curated dataset resulting from this approach, called CuREV, offers cleaner, higher-quality, and easier-to-learn-from comments that lead to measurable improvements in downstream automation tasks, namely review comment generation and code refinement. Building on this, we propose an improved pipeline, guided by high-quality exemplars, that enhances the realism and diversity of curated review comments. This method first separates the dataset into high-quality and low-quality reviews, based on a systematic quality assessment using an evaluation framework. High-quality comments are preserved in their original form and further used as in-context exemplars to inspire the reformulation of low-quality comments. By varying the exemplars provided, the reformulated comments are not only clearer and more actionable but also exhibit a broader range of writing styles, making them more realistic and human-like.

CommentsarXiv admin note: substantial text overlap with arXiv:2502.03425

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑