AI 中文总结
该研究针对现有 LLM 评审系统忽视论文内部主张与方法学证据不匹配的问题,提出 intra-paper claim verification 框架,经实验验证其生成的评审意见与人类评审关注点高度契合。
AI 中文摘要
科学投稿数量的不断增长,促使人们对使用大语言模型(LLM)辅助同行评审产生了兴趣。现有的自动化新颖性评估方法通常会将论文声称的贡献与现有文献进行比较,隐含假设这些贡献在论文本身中得到了准确实现。然而,人类评审者经常质疑新颖性主张,并非因为存在相似的想法,而是因为论文中呈现的方法学证据不足以支持这些主张。这种声称的贡献与方法学实现之间的内部 mismatch (不匹配),是当前基于 LLM 的评审系统很少研究的问题。为解决这一差距,我们提出了论文内部主张验证(intra-paper claim verification)框架,该框架用于评估论文中阐述的新颖性主张是否得到用于实现这些主张的方法的佐证。该框架利用 LLM 从引言中提取新颖性主张,检索与主张相关的方法学证据,并评估这些方法是否佐证了所述贡献。评估由从 182 篇 ICLR 2025 论文收集的人类同行评审中归纳得出的、受评审者启发的评估标准指导。这些标准涵盖了与新颖性、方法学、清晰度及其他问题相关的反复出现的评审者关注点,用于生成结构化的、类似评审者风格的主张佐证评估。我们通过将 LLM 生成的评审意见与已接收和已拒绝论文的平衡子集上的人类评审者关注点进行比较,对该框架进行评估。人类评估表明,框架生成的评估与人类评审者关注点之间存在显著的一致性,尤其是在与新颖性相关的问题上。BERTScore 进一步将对应的人类-LLM 评审对与不匹配的对照组区分开来,表明该框架捕获了与人类评审者观察一致的关注点。
英文摘要
The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing automated novelty assessment approaches typically compare a paper's claimed contributions against prior literature, implicitly assuming that these contributions are accurately realized in the work itself. Human reviewers, however, frequently challenge novelty claims not because similar ideas already exist, but because the methodological evidence presented in the paper does not adequately support them. This internal mismatch between claimed contributions and methodological realization is rarely examined by current LLM-based review systems. To address this gap, we introduce intra-paper claim verification, a framework that evaluates whether novelty claims articulated in a paper are substantiated by the methods used to realize them. The framework employs an LLM to extract novelty claims from the introduction, retrieve claim-relevant methodological evidence, and assess whether the methods substantiate the stated contributions. Assessment is guided by reviewer-inspired evaluation criteria derived inductively from human peer reviews collected from 182 ICLR 2025 papers. These criteria capture recurring reviewer concerns related to novelty, methodology, clarity, and other issues and are used to generate structured reviewer-style assessments of claim substantiation. We evaluate the framework by comparing LLM-generated review comments against human reviewer concerns on a balanced subset of accepted and rejected papers. Human evaluation demonstrates significant alignment between framework-generated assessments and human reviewer concerns, particularly for novelty-related issues. BERTScore further distinguishes corresponding human-LLM review pairs from mismatched controls, indicating that the framework captures concerns consistent with human reviewer observations.
Comments9 pages, 8 figures. Source code, prompts, evaluation materials, and supporting data available at https://github.com/Ranjitha2493/intra-paper-claim-method-verification