发表机构
Fudan University; The University of Tokyo; AtomInnoLab; Atom Infinite Pte. Ltd.(复旦大学; 东京大学; AtomInnoLab; Atom Infinite私人有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NovGauge是一个细粒度基准,通过三维标注和级联诊断流程评估18个LLM的论文新颖性判断,发现高幻觉率和证据不忠实问题,表明现有模型在此任务上仍不可靠。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于主要人工智能会议的同侪评审中,然而新颖性评估仍然是一个持续的薄弱环节。现有基准将新颖性作为单一整体分数进行评估,这使得难以诊断模型在哪个维度上判断失误,或者其证据是否忠实可靠。我们提出了NovGauge,一个以人类标注为锚点的细粒度新颖性评估诊断基准。该基准包含619对论文和50组多论文集合,数据来源于两个专家渠道:ICLR审稿人重叠声明和综述共被引关系。每个实例在三个维度上独立标注:任务、问题和方法,分别涵盖应用目标、技术挑战和解决方案途径。我们提出了一种级联诊断流程,用于验证各维度的正确性、证据依据和逻辑支撑。对18个LLM的评估显示,各维度的幻觉率在0%至39%之间,而在非幻觉的正确正向判断中,超过70%的引用证据无法在逻辑上支撑所述理由。表现最佳的模型GPT-5.5在验证后的F1分数上达到43%至72%,而大多数模型在忠实性验证后保留的原始F1分数不足一半。这些结果表明,当前LLM在科学新颖性评估方面仍远未达到可靠水平,尤其是在正确性依赖于忠实证据依据的情况下。
英文摘要
Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful. We present NovGauge, a human-anchored benchmark for fine-grained novelty assessment diagnosis. The benchmark contains 619 paper pairs and 50 multi-paper sets, drawn from two expert sources: ICLR reviewer overlap claims and survey co-citations. Instances are independently labeled along three dimensions: task, problem, and method, capturing application goals, technical challenges, and solution approaches. We propose a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support. Evaluation of 18 LLMs shows hallucination rates ranging from 0% to 39% across dimensions, and among non-hallucinated correct-positive judgments, over 70% cite evidence fails to logically support the stated reason. The best-performing model, GPT-5.5, achieves 43-72% Verified F1 across dimensions, while most models retain less than half of their raw F1 after faithfulness verification. These results suggest that current LLMs remain far from reliable scientific novelty assessment, particularly when correctness is conditioned on faithful evidence grounding.