arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

旧思想,新问题:基于LLM的新颖性评估的不稳定性

Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation

Noy Sternlicht, Simra Shahid, Peter Jansen, Daniel S. Weld, Pao Siangliulue, Tom Hope

arXiv 2610.02022首次发表:更新:

发表机构

Hebrew University of Jerusalem; Allen Institute for AI; Microsoft; University of Arizona; University of Washington(耶路撒冷希伯来大学; 艾伦人工智能研究所; 微软; 亚利桑那大学; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过系统性对照实验揭示,基于LLM的新颖性评估对提示设计极为敏感,微小的变化可大幅改变判定结果,现有评估方法不可靠,需开发稳健的新颖性评估方法。

AI 中文摘要

自动化构思系统通常根据其产生的想法的新颖性进行评估,而这一判断越来越多地交由大型语言模型负责。这类评判者通常是临时构建的,即便经过验证,也往往是在人类撰写的论文上进行验证,而非针对它们本应评分的生成想法。那么,新颖性评判者的表现如何?并不好。我们提出了一项关于新颖性评估设计选择的系统性对照研究。我们首先自动构建了一个评估集,从OpenReview中挖掘评审者明确肯定或质疑论文原创性的段落,并仅保留在其研究领域极端情况下达成一致同意的投稿;我们将这些段落与来自普通LLM生成器的想法配对。在六位评判者中,我们发现微小的提示设计选择会产生巨大的后果;例如,仅告知评判者评审者认为一个想法新颖而另一个不新颖,就能改变其超过一半的相同想法对上的判定,将成对准确率移动超过50个百分点,有时甚至将其推至低于随机水平。同样的改变对一位评判者有帮助,却对另一位有害。检索和更大的推理预算帮助甚微,而两个专门构建的新颖性评估器被我们最廉价的提示基线所超越。这些结果对自动化构思系统所报告的新颖性提升提出了质疑,并呼吁采用稳健的新颖性评估方法。

英文摘要

Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant to score. So, how do novelty judges perform? Not well. We present a systematic controlled study of novelty evaluation design choices. We first build an evaluation set automatically, mining OpenReview for passages where reviewers explicitly affirm or dispute a paper's originality and keeping only submissions with unanimous agreement at the extremes of their research area; we pair these with ideas from a vanilla LLM generator. Across six judges, we find that small prompt design choices have large consequences; e.g., simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical idea pairs it is shown, shifting pairwise accuracy by over 50 points and occasionally pushing it below chance. The same change helps one judge and hurts another. Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about reported novelty gains of automated ideation systems, and call for robust novelty evaluation methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑