发表机构
University College London; Amazon; Zhejiang University(伦敦大学学院; 亚马逊; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型语言模型生成的看似合理但无效的代码审查评论,提出CRJudgeBench基准和基于仓库的智能体评判器Sentinel,通过迭代动作级学习显著提升技术可信度判断准确率。
AI 中文摘要
大型语言模型能够生成看似合理的代码审查评论,但此类评论可能包含技术上不正确的说法,从而误导开发者。我们研究技术可信度判断:确定审查评论的核心技术主张是否正确,并且是否适用于所审查代码在其仓库上下文中的情况。现有的代码审查基准主要评估审查生成、问题发现或一般评论质量,但并未直接评估智能体能否判断单个审查评论的技术可信度。为填补这一空白,我们引入了CRJudgeBench,一个由真实拉取请求和专家验证的扰动构建的1199个实例的基准,涵盖可信赖和看似合理但不可信赖的评论。我们进一步提出了Sentinel,一个基于仓库的智能体评判器,在做出判断之前主动收集代码证据以验证审查评论。从Qwen3-Coder-30B-A3B-Instruct出发,Sentinel通过在CRJudgeBench训练集上进行基于特权教师的迭代动作级学习进行训练。在359个实例的CRJudgeBench测试集上,Sentinel达到了76.60%的准确率,比GLM-5.3高出6.13个百分点,比其基础模型高出19.78个百分点。这些结果表明,即使是最先进的一般用途大型语言模型也难以识别不可信赖的评论,而迭代动作级学习显著提高了基于仓库的可信度判断的准确性。我们的数据集可在以下网址获取:此https URL。
英文摘要
Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible but untrustworthy comments. We further present Sentinel, a repository-grounded agentic judge that actively gathers code evidence to verify review comments before making judgments. Starting from Qwen3-Coder-30B-A3B-Instruct, Sentinel is trained on the CRJudgeBench training split through iterative action-level learning from a privileged teacher. On the 359-instance CRJudgeBench test set, Sentinel achieves 76.60\% accuracy, outperforming GLM-5.3 by 6.13 percentage points and its base model by 19.78 points. These results show that even state-of-the-art general-purpose LLMs struggle to identify untrustworthy comments, while iterative action-level learning substantially improves the accuracy of repository-grounded trustworthiness judgments. Our dataset is available at https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark
Comments26 pages, 5 figures, and 14 tables. Under review at ICLR 2027. Dataset available at https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark