发表机构
Technion; Sakana AI; FARS(以色列理工学院; Sakana AI; FARS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对AI科学家系统的评估难题,提出基于多模型LLM的自动化评审基准,发现FARS基准论文表现最优,Gemini与Claude评估一致性强,GPT-5.4标准不同,建立了首个定量基准。
AI 中文摘要
具备自主研究能力的AI科学家系统有望大幅加速科学发现,但评估和比较AI生成论文的质量仍是未解决的挑战。我们提出并实施了一套严格的基准测试协议,利用前沿大语言模型构建自动化同行评审系统,从原创性、科学严谨性、清晰度和重要性四个核心维度评估科学论文。我们评估了四个领先的AI科学家框架:Sakana AI(v1和v2)、CycleResearcher和Data-to-Paper。每个框架均在某商业自主AI科学家公司FARS发布的15组一致研究提案上运行,共生成60篇论文,我们将这些论文与15篇FARS基准论文一同评估。使用三个独立的LLM评审员(GPT-5.4、Gemini和Claude),我们发现FARS基准论文显著优于所有对比框架,在1-5分制下取得2.14至2.47的平均分,而其他系统的分数为1.00至1.87。值得注意的是,在Gemini和Claude的评估中,FARS的分数比次优系统高出2倍以上。我们发现Gemini和Claude之间存在强一致性(ρ=0.907,p<0.001),且两者与综合分数的相关性极强(ρ=0.961,p<0.001),验证了自动化评估的可靠性。然而,GPT-5.4的一致性较弱(ρ≈0.32),表明其使用不同的标准评估论文。这些结果建立了首个针对AI科学家系统的定量基准,证明多模型LLM评估为评估自主研究质量提供了可扩展、一致的框架。
英文摘要
AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose and implement a rigorous benchmarking protocol using an automated peer-review system that harnesses frontier large language models to assess scientific papers across four core dimensions: originality, scientific rigor, clarity, and significance. We evaluate four leading AI Scientist frameworks: \textit{Sakana AI (v1 & v2)}, \textit{CycleResearcher}, and \textit{Data-to-Paper}. Each framework was run on a consistent set of 15 research proposals published by a commercial autonomous AI scientist company (FARS), generating 60 papers that we evaluate alongside 15 FARS benchmark papers. Using three independent LLM reviewers (GPT-5.4, Gemini, and Claude), we find that FARS benchmark papers significantly outperform all competing frameworks, achieving mean scores of 2.14--2.47 on a 1--5 scale compared to 1.00--1.87 for other systems. Notably, FARS scores are more than 2$\times$ higher than the next-best systems on Gemini and Claude evaluations. We find strong agreement among Gemini and Claude ($ρ$ = 0.907, $p < 0.001$), and both correlate extremely strongly with the synthesis score ($ρ$ = 0.961, $p < 0.001$), validating the reliability of automated evaluation. However, GPT-5.4 exhibits weaker agreement ($ρ\approx 0.32$), suggesting it evaluates papers using different criteria. These results establish the first quantitative benchmark for AI Scientist systems and demonstrate that multi-model LLM evaluation provides a scalable, consistent framework for assessing autonomous research quality.