更多批评并不意味着更好的评审:EquiReview-R
More Criticism Does Not Make a Better Review: EquiReview-R
查看机构详情
- College of Systems Engineering, National University of Defense Technology(国防科技大学系统工程学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对AI评审中更多批评未必更好的问题,提出EquiReview-R模型,基于证据优化结构化关切集,在论文评审任务中降低过度批评比例,发布证据关联语料库ReviewTrace用于相关研究。
中文摘要 AI 辅助
AI评审如今能生成许多具体批评,但更多批评未必是更好的评审。评审可能遗漏关键缺陷,或保留无现有证据支持的指控。这些失败需要相反的修正,然而面向生成的系统和聚合指标模糊了这种区别。因此,我们将AI辅助评审重新定义为基于证据的结构化关切集优化,将遗漏和过度批评视为两种独立风险。基于该框架,我们提出EquiReview-R,它针对局部证据解决现有关切,从独立和评审条件视角搜索缺失问题,并返回“停止”“继续”或“弃权(不执行)”。为揭示驱动该设计的失败模式,我们构建了证据关联轨迹语料库。其回顾分析表明,修正必须先于进一步搜索:高召回评审中几乎所有关切都缺乏明确证据处置,而早期优化机制无法修正它们。在一组先前未见过的论文的冻结队列上,EquiReview-R满足主要遗漏的预设非劣性准则,将主要过度批评从15.5%降至8.1%,单侧遗漏上界达9.9%,同时对52.4%的论文“停止”。计算匹配对照、受控配对和 ablation 实验显示,该收益来自优化而非额外推理或更短输出。我们发布该语料库作为ReviewTrace,这是用于研究评审优化、分歧和溯源的证据关联资源。
英文摘要
AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet generation-oriented systems and aggregate measures obscure the distinction. We therefore recast AI-assisted review as evidence-guided refinement of a structured concern set, with omission and overcritique treated as separate risks. Building on this formulation, we introduce EquiReview-R, which resolves existing concerns against localized evidence, searches for missing issues from independent and review-conditioned perspectives, and returns stop, continue, or defer. To expose the failure mode that motivates this design, we construct an evidence-linked trajectory corpus. Its retrospective analysis shows why revision must precede further search: nearly all concerns in a high-recall review lack a definitive evidential disposition, while an earlier refinement mechanism cannot revise them. On a frozen cohort of previously unseen papers, EquiReview-R satisfies the prespecified non-inferiority criterion for major omission, reduces major overcritique from 15.5% to 8.1%, and attains a one-sided omission upper bound of 9.9% while stopping on 52.4% of papers. Computation-matched controls, controlled pairs, and ablations show that the gain comes from revision rather than extra inference or shorter output. We release the corpus as ReviewTrace, an evidence-linked resource for studying review revision, disagreement, and provenance.