arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10511cs.CL

清晰的直觉还是真正的分析?评估大语言模型作为评审员的同行评审中的认知可靠性基准

Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Nuo Chen, Qian Wang, Qingyun Zou, Bingsheng He

AI总结:

研究大语言模型评审员与人类评审员对同行评审评估的差异,将卡尼曼双过程理论转化为评分标准并发布Kahneman4Review基准,通过多方面发现揭示问题,强调可靠基准应区分分析形式与认知功能并给出设计选择。

AI中文摘要:

当一个大语言模型评审员称一篇同行评审具有分析性,而一个人类委员会称另一篇评审质量高时,他们追踪的是同一件事吗?我们认为并非如此,且这种差异在哲学上很重要。我们将卡尼曼的双过程理论转化为同行评审的结构化评分标准,并发布了Kahneman4Review,这是一个包含3563篇有评分的评审的基准,根据九个理论驱动的文本维度、八个偏差诊断和一个连续的推理质量分数进行评分。有三个关于可信度的发现:决策层级与评分标准中基于文本的认知质量代理没有明显对齐;公开展示的代理评审比汇总的人类评审获得更高的原始分数,但长度和发表场所解释了大部分差距,且样本不是论文配对的;ICLR评审文本诊断在2022 - 2023年过渡时发生变化,与大语言模型的广泛应用时间一致,但未确定其原因。一个匹配的功能探测试点进一步表明,该评分标准能区分旨在对比真正的故障发现与表面流畅性的文本探测。我们认为,一个值得信赖的大语言模型评审员可靠性基准必须将分析形式与认知功能分开,并朝着该目标提出具体的设计选择。可通过此https URL获得交互式演示。

英文摘要:

When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the difference matters philosophically. We operationalise Kahneman's dual-process theory into a structured rubric for peer review and release Kahneman4Review, a benchmark of 3,563 rated reviews scored along nine theoretically motivated textual dimensions, eight bias diagnostics, and a continuous reasoning-quality score. Three findings bear on trustworthiness: decision tier is not detectably aligned with the rubric's text-grounded epistemic-quality proxy; public-showcase agentic reviews receive higher raw scores than pooled human reviews, but length and venue explain most of the gap and the samples are not paper-paired; and ICLR review-text diagnostics shift at the 2022--2023 transition, temporally coincident with widespread LLM availability but without identifying its cause. A matched function-probe pilot further shows that the rubric distinguishes textual probes designed to contrast genuine fault-finding with surface fluency. We argue that a trustworthy reliability benchmark for LLM judges must separate analytical form from epistemic function, and propose concrete design choices toward that goal. An interactive demo is available at https://huggingface.co/spaces/nuojohnchen/Kahneman4Review.

补充信息

↑