发表机构
Fudan University; Zhongguancun Academy; Tsinghua University; Zhongguancun Institute of Artificial Intelligence(复旦大学; 中关村学院; 清华大学; 中关村人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Ideation Arena平台,通过105名计算机科学研究者的双盲成对比较构建LLM生成研究创意的Elo评级排行榜,发现部分智能体框架可提升创意质量,而现有LLM评判者难以复现专家偏好。
AI 中文摘要
评估大语言模型(LLM)生成的研究创意颇具难度,因为其科学价值无法通过客观标准完全判定,且不存在单一参考答案来界定何为优质创意。为应对这一挑战,我们推出了Ideation Arena,这是一个采用对战式人类评估的平台,用于评估研究创意。Ideation Arena评估了14个前沿LLM以及基于2个基础模型构建的5个研究智能体架构所生成的创意。为确保共同起点,Ideation Arena从参与研究人员熟悉的论文中构建共享文献上下文,并将相同上下文提供给所有LLM和智能体。我们从105名活跃的计算机科学研究人员处收集了超过6000个双盲成对比较结果,构建了在共享封闭上下文协议下计算机科学提案阶段专家偏好的Elo评级排行榜。我们通过评分者间一致性和稳健性分析验证了该排行榜,结果表明,在标注者组成和领域覆盖发生变化时,排行榜保持稳定。我们的结果显示智能体有效性存在显著差异:部分框架提升了其基础模型的创意质量,而另一些则几乎没有益处,甚至表现逊于其基础模型。我们进一步构建了Ideation Arena Eval,这是一个用于评估自动评估器是否与研究创意领域的人类偏好对齐的基准。对当前LLM评判者的实验表明,它们仍无法可靠复现专家偏好,最佳评判者在整体质量上达到72.56%的软准确率。我们的代码、数据和排行榜可在该https URL获取。
英文摘要
Evaluating research ideas generated by LLMs is difficult because their scientific value cannot be fully determined by objective criteria, and no single reference answer specifies what counts as a good idea. To address this challenge, we introduce Ideation Arena, a battle style platform that evaluates research ideas through pairwise human assessment. Ideation Arena evaluates ideas generated by 14 frontier LLMs and 5 research agent architectures built on 2 base models. To ensure a common starting point, Ideation Arena builds shared literature contexts from papers familiar to the participating researchers and provides the same contexts to all LLMs and agents. We collect over 6,000 double blind pairwise comparisons from 105 active computer science researchers and construct an Elo rating leaderboard of proposal-stage expert preferences in computer science under a shared closed-context protocol. We validate the rankings through interrater agreement and robustness analyses, showing that the leaderboard remains stable under changes in annotator composition and domain coverage. Our results show substantial variation in agent effectiveness, with some frameworks improving ideation quality over their backbones and others offering little benefit or even underperforming their base models. We further construct Ideation Arena Eval, a benchmark for assessing whether automated evaluators align with human preferences in research ideation. Experiments with current LLM judges show that they still cannot reliably reproduce expert preferences, with the best judge reaching 72.56% Soft Accuracy on Overall Quality. Our code, data, and leaderboards are available at https://github.com/foss12138/Research-Ideation-Arena.