CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?
CRJudgeBench:AI能否检测看似合理但无效的代码审查?
机构 * University College London(伦敦大学学院) ; Amazon(亚马逊) ; Zhejiang University(浙江大学)
AI总结 针对大型语言模型生成的看似合理但无效的代码审查评论,提出CRJudgeBench基准和基于仓库的智能体评判器Sentinel,通过迭代动作级学习显著提升技术可信度判断准确率。
Comments 26 pages, 5 figures, and 14 tables. Under review at ICLR 2027. Dataset available at this https URL (https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark)