arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39027cs.CL

可信 AI 审稿人缺失的一环:从修辞鲁棒性基准测试到 SciCore 审稿

A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

Chenguang Wang, Ming Li, Chengrui Fan, Jianpeng Chen, Han Chen, Tianyi Zhou, Dawei Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

针对 AI 审稿人对同科学内容不同措辞判断不一的问题,提出修辞鲁棒性定义、RobustReview 基准(1260 版本)及 SciCore 双分支审稿模型,在 GPT-5.5 对比中实现领先的稳定性-区分度平衡。

中文摘要 AI 辅助

AI 审稿人可能会对报告相同科学内容但措辞不同的稿件给出不同的判断,从而可能奖励修辞优化而非科学改进。我们将修辞鲁棒性(Rhetorical Robustness)定义为跨内容保持改写(content-preserving rewrites)的稳定性与跨论文的区分度(discrimination)的联合要求。我们引入了 RobustReview,一个受控的全稿件基准测试,包含 1,260 个稿件版本,并评估了 30 种审稿人配置。该基准测试揭示了虚假鲁棒性(false robustness),即低改写敏感性与跨论文分数坍缩(score collapse)同时出现,并表明人类对齐(human alignment)与修辞鲁棒性对审稿人的排序不同。此外,所评估的以内容为中心的提示协议(content-focused prompting protocol)并未在不同骨干模型(backbones)上持续提升鲁棒性。受这些发现的启发,我们提出了 SciCore,一种双分支审稿人,它将全稿件判断与基于提取的结构化科学核心(science core)的判断进行平均。这种设计将稿件级评估与内容归一化视图相结合,旨在降低修辞敏感性。在我们主要的 GPT-5.5 对比中,SciCore 在基准测试的审稿人中取得了领先的联合稳定性-区分度(joint stability-discrimination)表现,同时保持了具有竞争力的人类对齐。这些结果将修辞鲁棒性确立为一个独立的评估目标,并展示了科学核心审稿(science-core review)在改进这一目标上的潜力。

英文摘要

AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently. Moreover, the evaluated content-focused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscript-level assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.

发表机构

  • Virginia Tech(弗吉尼亚理工大学)
  • University of Maryland(马里兰大学)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑