发表机构
University of California, Santa Cruz(加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过合成宇宙基准检验AI科学家能否区分预测充分的定律与恢复机制的定律,发现两者可分离,需分别测试。
AI 中文摘要
一个科学智能体能否区分它从证据中推断出的定律与它仅仅识别的定律?我们引入了合成宇宙(Synthetic Universes),这是一个受控基准,将著名的经典世界与由邻近非经典机制驱动的匹配扭曲孪生世界配对。我们对每个报告的定律进行两次评估:一是在留出的延续和迁移设置上执行它,二是独立检查它是否恢复了生成机制。在预先指定的60单元研究的当前检查点中,22次试验已评分,另有1次运行因基础设施故障而终止。在20次孪生试验中,8次通过预测验证,而5次恢复了生成器。这种分离是双向的:6个可解析输出成功预测但遗漏了机制,而3个恢复了机制但未能通过预测展开。Drag表现出第一种模式(5/5预测通过,1/5机制恢复);Gravity表现出第二种模式(1/5预测通过,4/5机制恢复)。由于匹配的著名对照、校正的可识别性扫描和证据阶梯(Evidence Ladder)仍未完成,我们不声称存在确证性的因果先验冲突效应。相反,已完成的运行确立了一个更窄的验证结果:预测充分性和机制恢复是不同的科学主张,需要不同的测试。
英文摘要
Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms. We evaluate each reported law twice: by executing it on held-out continuations and transfer settings, and by independently checking whether it recovers the generating mechanism. In the current checkpoint of a pre-specified 60-cell study, 22 trials were graded and one additional run ended in infrastructure failure. Among 20 twin trials, 8 pass predictive verification while 5 recover the generator. The dissociation is bidirectional: six parsable outputs predict successfully while missing the mechanism, whereas three recover the mechanism but fail predictive rollout. Drag exhibits the first pattern (5/5 predictive pass, 1/5 mechanism recovery); Gravity exhibits the second (1/5 predictive pass, 4/5 mechanism recovery). Because matched famous controls, the corrected identifiability sweep, and the Evidence Ladder remain incomplete, we do not claim a confirmatory causal prior-conflict effect. Instead, the completed runs establish a narrower verification result: predictive adequacy and mechanism recovery are distinct scientific claims and require distinct tests.