发表机构
Zhejiang University; Nanyang Technological University; Sun Yat-sen University; Southeast University; Zhejiang Normal University; The Hong Kong University of Science and Technology(浙江大学; 南洋理工大学; 中山大学; 东南大学; 浙江师范大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一个过程中心基准,通过比较Gold与Predicted过程变量的决定价值,诊断AI辅助同行评审中最终决定是否得到充分评审证据支持,实验表明存在稳定差距。
AI 中文摘要
同行评审是科学质量控制的核心。然而,现有对AI辅助同行评审的评估主要关注生成评审的整体质量或最终决定的准确性。因此,它们提供的证据有限,无法证明模型决定是否得到充分且可靠的评审证据支持。我们引入了一个面向AI辅助同行评审的过程中心诊断基准。它使用(x,$z_s$,$z_c$,$z_r$,y)来表示论文内容、摘要、批评、建议和决定。我们将来自PeerRead、NLPeer ARR-22和OpenReview-ICLR的异构评审记录转换为过程对齐数据。我们的基准使用从论文内容直接进行决定预测(Direct)作为基线。它比较了Gold过程变量和Predicted过程变量的决定价值,并进行了阶段级评估、链一致性评估和干预敏感性分析。在三个数据集和六个模型上的实验表明,Gold过程变量通常具有更高的决定价值。对于主要分析模型,Gold-Predicted差距在不同数据集和随机种子间保持稳定。这一差距在大多数模型-数据集组合中也得到重现。尽管模型生成的中间评审文本在相邻阶段间表现出相对较高的局部一致性,但最终决定并未得到先前评审证据的一致支持。我们的基准针对旨在辅助而非取代人类评审者的AI系统。它为评估其评审过程的可靠性提供了一个透明且可审计的诊断工具。
英文摘要
Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of generated reviews or the accuracy of final decisions. They therefore provide limited evidence about whether model decisions are supported by sufficient and reliable review evidence. We introduce a process-centric diagnostic benchmark for AI-assisted peer review. It uses (x,$z_s$,$z_c$,$z_r$,y) to represent the paper content, summary, critique, suggestion, and decision. We convert heterogeneous review records from PeerRead, NLPeer ARR-22, and OpenReview-ICLR into process-aligned data. Our benchmark uses direct decision prediction from the paper content (Direct) as its baseline. It compares the decision value of Gold-process variables and Predicted-process variables, and conducts stage-level evaluation, chain-consistency evaluation, and interventional sensitivity analysis. Experiments across three datasets and six models show that Gold-process variables generally have higher decision value. For the main analysis model, the Gold--Predicted gap remains stable across datasets and random seeds. This gap is also reproduced in most model--dataset combinations. Although model-generated intermediate review texts show relatively high local consistency across adjacent stages, the final decisions are not consistently supported by the preceding review evidence. Our benchmark targets AI systems designed to assist rather than replace human reviewers. It provides a transparent and auditable diagnostic tool for evaluating the reliability of their review processes.