发表机构
University of California San Diego; University of California, Irvine; Cushing Academy; University of California, Santa Barbara(加州大学圣地亚哥分校; 加州大学欧文分校; 卡欣学院; 加州大学圣巴巴拉分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出科学定律发现基准SCILAWS-BENCH,通过真实与平行双设置评估LLM的科学定律发现能力,发现预测拟合与科学有效性背离、记忆影响公式复现及存在选择瓶颈等现象,为相关评估提供新视角。
AI 中文摘要
科学定律发现长期以来是科学进步的核心,通过在科学约束下进行假设生成、观测检验与迭代完善的循环推进。随着大型语言模型(LLM)能力的提升及其在科学人工智能(AI for Science)中的作用扩大,它们能否真正发现科学定律、该能力应如何评估仍是未解决的问题。然而现有评估常通过合成设置简化发现过程,或复用可能已为LLM所熟悉的已发表目标。因此,我们推出SCILAWS-BENCH——基于已发表研究与真实科学数据构建的科学定律发现基准,它包含来自381篇科学论文的118个问题,覆盖6个科学领域的291条候选定律和约800万条真实数据点。每个问题以两种互补设置实现:(1)SCILAWS-REAL要求模型从固定真实观测中提出定律,评估保留的预测拟合度与源自源文献的科学有效性;(2)SCILAWS-PARALLEL要求模型主动查询经残差校准的世界,恢复源自已发表形式的合成隐藏定律。这种双设置任务设计保留了每个问题的科学背景,同时分别评估固定记录的定律发现与主动恢复新合成隐藏定律的能力。我们发现预测拟合度可能与科学有效性背离,记忆会影响模型是否复现或超越已发表公式,且我们的N选一研究揭示了选择瓶颈。本研究为评估用于科学发现的AI提供了基于论文的基准与新的实证视角。项目页面:this https URL
英文摘要
Scientific law discovery has long been central to scientific progress, proceeding through iterative cycles of generating hypotheses, testing them against empirical evidence, and refining them under scientific constraints. As large language models (LLMs) become increasingly involved in scientific research, whether they can discover scientific laws and how to evaluate this ability remain open questions. A central evaluation challenge is to move beyond familiar published equations while keeping discovery tasks grounded in scientific data and constraints. We introduce SciLaws-Bench, a curated collection of scientific task packages grounded in the source literature, each linking a scientific problem, supporting data, published reference equations, and scientific-validity rubrics. Through agent-assisted curation and human verification, we assemble 118 problems spanning six disciplines, drawing on 381 papers, 291 candidate laws, and roughly 8M data points. Each problem supports two complementary evaluation settings. SciLaws-Real uses fixed scientific data to evaluate proposed laws for held-out predictive fit and scientific validity. SciLaws-Parallel evaluates recovery of a newly synthesized structural variant of a published equation through active queries to a simulator calibrated to the source data. Our evaluation reveals three limitations: good predictive fit need not imply scientific validity, recovering a published formula does not establish recovery of its new structural terms, and candidate selection remains a bottleneck in scientific law discovery. Project page: https://yiyihum.github.io/SciLaws-Bench
Comments41 pages, 14 figures