arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以貌取评:基于LLM的同行评审评估指标的可靠性分析

Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics

Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri

arXiv 2609.23264首次发表:更新:

发表机构

University of California, Berkeley; University of Toronto(加州大学伯克利分校; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出统计框架,通过对比原始人工评审与保留语义的LLM改写,发现29个同行评审评估指标中23个对改写敏感,表明多数指标将评审质量与语言呈现混淆,需验证稳健性。

AI 中文摘要

同行评审评估正越来越多地使用LLM-as-a-judge指标进行自动化,但这带来了测量风险。一篇评审可能因其流畅、有条理且措辞精炼而获得高分,而非因其对论文提供了强有力的评估。这一风险在AI辅助评审中尤为重要,因为评审者可能使用LLM来改进清晰度或呈现方式,同时保留其潜在的判断。我们提出了一个统计框架,用于测试同行评审评估指标是否捕捉了超越表面语言形式的实质性评审质量。该框架将原始人工评审与忠实的LLM改写版本进行比较,这些改写版本在改变措辞和呈现方式的同时保留了相同的评估内容。利用一个包含来自ICLR和NeurIPS的674篇人工评审的4,044个保留语义的改写数据集,我们通过表面敏感性和稳健性的互补测试,评估了来自四项先前工作的29个面向内容的同行评审评估指标。尽管这些指标旨在捕捉超越表面层面、依赖写作特征的评审属性,我们发现对改写的敏感性普遍存在。在我们的主要分析中,23个指标对评估内容保持不变的评审给出了显著不同的分数,而只有6个满足我们的稳健性标准。这些模式在两个LLM评判模型之间基本一致,表明该问题并非特定于单个评判者。这些发现表明,许多同行评审评估指标部分地将评审质量与语言呈现混淆,并表明在将这些指标用于比较人工撰写、AI辅助和AI生成的评审之前,应验证其对保留语义改写的稳健性。

英文摘要

Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted reviewing, where reviewers may use LLMs to improve clarity or presentation while preserving the underlying judgments. We propose a statistical framework for testing whether peer-review evaluation metrics capture substantive review quality beyond surface-level linguistic form. The framework compares original human reviews with faithful LLM rewrites that preserve the same evaluative content while changing wording and presentation. Using a dataset comprising 4,044 meaning-preserving rewrites derived from 674 human reviews from ICLR and NeurIPS, we evaluate 29 content-oriented peer-review evaluation metrics drawn from four prior works through complementary tests of surface sensitivity and robustness. Although these metrics are intended to capture review properties beyond surface-level, writing-dependent characteristics, we find that sensitivity to rewriting is widespread. Under our primary analysis, 23 metrics assign significantly different scores to reviews whose evaluative content is preserved, while only six satisfy our robustness criterion. The patterns are largely consistent across two LLM judge models, suggesting that the issue is not specific to a single judge. These findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.

CommentsAccepted at CIKM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑