arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05788cs.AI

超越模仿审稿人:评估用于投稿前同行评审的大语言模型

More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review

  • University of Minnesota(明尼苏达大学)
  • Invariant Tech Inc.(Invariant Tech 公司)

机构由 AI 辅助整理,请以论文原文为准。

Pouya Parsa, Amin Rezaei

AI总结:

本研究评估面向作者的LLM投稿前评审系统,发现大型候选池覆盖率高但压缩效果差,代表性选择和匹配器敏感性是主要瓶颈。

AI中文摘要:

同行评审反馈往往来得太晚,作者无法进行有意义的修改。我们研究了一个面向作者的LLM系统,该系统将部分压力测试提前到投稿之前:它生成大量原子化关注点,并将其压缩成一份简短报告。我们评估了该系统与历史评审的一致性,并分别评估了其遗漏关注点的可能有效性。从10,000篇ICLR 2026投稿中,我们使用了3,398篇具有评审前可访问版本的稿件。在十篇论文的诊断测试中,独立采样覆盖了44.9%的历史问题;去重和补充填充达到了78.7%的严格覆盖率和84.9%的严重性加权覆盖率,但请求量增加了3.6倍,令牌数增加了5.2倍。一个隐藏的Top-32预言机保留了256个候选池的全部79.3%加权覆盖率,而仅基于论文的选择器仅保留了40%至44%。因此,LLM评审通过大型候选池提供了广泛的覆盖率,但压缩效果不佳;消融实验表明,代表性选择和匹配器敏感性是造成这一差距的主要来源。

英文摘要:

Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We evaluate agreement with historical reviews and, separately, the possible validity of concerns they omit. From 10,000 ICLR 2026 submissions, we use 3,398 manuscripts with accessible versions that predate review. On a ten-paper diagnostic, independent sampling covers 44.9% of historical issues; deduplication and refill reaches 78.7% strict and 84.9% seriousness-weighted coverage, at 3.6$\times$ more requests and 5.2$\times$ more tokens. A hidden Top-32 Oracle preserves the full 79.3% weighted coverage of a 256-candidate pool, but paper-only selectors retain only 40--44%. LLM review therefore provides broad coverage with a large candidate pool but compresses poorly; ablations identify representative selection and matcher sensitivity as the main sources of this gap.

补充信息

↑