arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考用于生成代码审查评论的训练数据

Rethinking Training Data for Generating Code Review Comments

Leonardo Centellas-Claros, Estefania Pakarati-Cofre, Juan Pablo Sandoval Alcocer, Diego Elias Costa

arXiv 2607.25851首次发表:更新:

AI 中文总结

研究代码审查评论生成,指出尽管有进展但生成的评论存在问题,通过实证检查识别出不匹配训练对并得出分类法,探讨纳入分类法能否改善问题,发现仍有挑战,认为改进需更多举措而非仅数据集清理。

AI 中文摘要

生成代码审查评论已成为自动代码审查中的一个突出研究方向,通常被表述为基于差异-评论对的文本生成任务。尽管基于学习的方法取得了进展,但生成的审查评论往往很通用、缺乏依据或不可操作。最近的研究表明审查评论数据集包含有噪声或不合适的训练实例,这促使基于大语言模型(LLM)的数据集清理方法出现。本文认为有问题的训练实例并非同质,且一些局限性源于任务表述本身更深层次的问题。通过对一个广泛使用的审查评论数据集的实证检查,我们识别出了不匹配的训练对:即代码更改与审查评论之间的关系无法为从局部输入生成可操作的审查反馈提供可靠学习信号的实例。我们得出了一种不匹配分类法,涵盖三个反复出现的来源:语义模糊、缺乏可操作性和上下文依赖性。我们进一步探讨将这种分类法纳入基于LLM的过滤是否能改善对有问题训练实例的识别,发现检测不匹配的训练对仍然具有挑战性。基于这些观察结果,我们认为改进审查评论生成仅靠数据集清理是不够的,还需要明确的有效性标准、更丰富的上下文输入以及与审查意图和可操作性一致的评估实践。

英文摘要

Generating code review comments has become a prominent research direction in automated code review, commonly formulated as a text generation task over diff-comment pairs. Despite advances in learning-based approaches, generated review comments are often generic, weakly grounded, or non-actionable. Recent studies have also shown that review comment datasets contain noisy or unsuitable training instances, motivating LLM-based dataset cleaning approaches. In this paper, we argue that problematic training instances are not homogeneous and that some limitations stem from deeper issues in the task formulation itself. Through an empirical inspection of a widely used review comment dataset, we identify misaligned training pairs: instances where the relationship between the code change and the review comment does not provide a reliable learning signal for generating actionable review feedback from localized inputs. We derive a taxonomy of misalignment capturing three recurring sources: semantic ambiguity, lack of actionability, and context dependence. We further explore whether incorporating this taxonomy into LLM-based filtering improves the identification of problematic training instances, observing that detecting misaligned training pairs remains challenging. Based on these observations, we argue that improving review comment generation requires more than dataset cleaning alone, motivating explicit validity criteria, richer contextual inputs, and evaluation practices aligned with review intent and actionability.

CommentsPaper accepted to ICSME 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑