发表机构
Carnegie Mellon University; Stanford University; Google DeepMind; Yale University; Massachusetts Institute of Technology; Columbia University; Princeton University; Flatiron Institute; Engram(卡内基梅隆大学; 斯坦福大学; 谷歌DeepMind; 耶鲁大学; 麻省理工学院; 哥伦比亚大学; 普林斯顿大学; 熨斗研究所; Engram)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EurekaBench是一个跨领域基准,通过26个长期任务和306个科学见解,评估AI智能体在科学发现中遵循约束、预测准确性和产生见解的能力,发现现有智能体过度优化预测而缺乏见解。
AI 中文摘要
当艾萨克·牛顿发现万有引力定律时,他通过一个迭代过程实现这一发现:分析观测数据(如行星模式),通过用数学方程描述模式来找到潜在机制,并根据月球轨道修正他的理论,从而揭示了一个惊人的见解:同一种力既支配着落下的苹果,也支配着绕行的行星。人工智能智能体是否可能做出类似的发现?为了衡量这种能力,我们引入了EurekaBench,一个跨领域的基准测试,用于测试AI智能体进行长期实验并发现解释观测结果的机制的能力。我们通过从这些机制中能够推导出的科学见解来评估它们。EurekaBench包含一个由专家验证的26个长期任务集合,涵盖神经科学、计算机科学、化学、天体物理学、地球物理学和等离子体物理学,总共有306个科学见解,这些发现机制预期应支持这些见解。我们的评估框架测试科学发现的三个维度:智能体遵循已知科学约束的能力、所发现机制的预测准确性,以及这些机制是否产生科学见解或为未来研究提供信息。我们的结果表明,当前的AI智能体往往过度专注于预测准确性优化,超越了人类科学家,但在推导科学见解方面则明显不足。
英文摘要
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.