arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14057cs.CL

衡量前沿大语言模型在自动化研究中的创造力

Measuring the Creativity of Frontier LLMs in Automated Research

Yiheng Zhao, Mengzhuo Chen, Chengming Hu, Pengyi Liao, Yihan Huang, Yiran Pang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一套衡量指标,从价值性和新颖性两个维度评估前沿大语言模型在自动化研究中的创造力,并发现变量级论文级新颖性最能预测研究表现。

中文摘要 AI 辅助

前沿大语言模型(LLMs)在自动化研究方面的能力日益增强,然而它们在此情境下的创造力尚未得到系统评估。本文提出了一套衡量指标,从价值性和新颖性两个维度评估创造力。价值性评估每个提出想法是否有用,而新颖性则从三个角度进行评估:同一想法是否曾出现过(精确匹配的论文级新颖性,Exact-Match P-Novelty);是否探索了先前未涉及的变量或变量组合(变量级论文级新颖性,Variable-level P-Novelty);以及该想法是直接遵循检索到的外部知识还是偏离了它(假设级新颖性,H-Novelty)。我们的评估显示,模型在大多数创造力指标上得分相对接近,但在变量级论文级新颖性上差异显著,这反映了研究空间探索的广度。进一步的相关性和想法级表现分析表明,变量级论文级新颖性是与研究表现关联最紧密的创造力维度。

英文摘要

Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. We propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives: whether the same idea has appeared before (Exact-Match P-Novelty), whether the modified variable or variable combination has been explored before (Variable-level P-Novelty), which reflects the breadth of research-space exploration, and whether the proposed idea is explicitly attributed to external knowledge in the model's reasoning (H-Novelty). Our evaluation shows that the models achieve relatively similar Valueness and Exact-Match P-Novelty scores, while differing substantially in Variable-level P-Novelty. H-Novelty is also consistently high among the models for which it can be evaluated. Notably, further correlation and idea-level performance analyses reveal a strong positive correlation between Variable-level P-Novelty and research performance.

发表机构

  • Concordia University(康考迪亚大学)
  • McGill University(麦吉尔大学)
  • Florida Atlantic University(佛罗里达大西洋大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑