学习超越理论:在初始假设空间之外的实验发现
Learning to Outgrow a Theory: Experimental Discovery Beyond the Initial Hypothesis Space
- Nanyang Technological University(南洋理工大学)
- Carnegie Mellon University(卡内基梅隆大学)
- Hunan University(湖南大学)
- Beijing Normal University(北京师范大学)
- University of Science and Technology of China(中国科学技术大学)
- University of the Chinese Academy of Sciences(中国科学院大学)
- Zhejiang Lab(之江实验室)
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出一种联合结构编辑与诊断实验的发现策略,通过类级可区分性和随时有效证据,在固定假设空间外实现模型类修订,显著提升精确恢复率并跨任务迁移。
AI中文摘要:
科学发现系统通常在固定的假设空间内优化实验。当所有可用候选都遗漏了相同的缺失机制时,这会产生一种失败模式:即使模型类系统性错误,候选之间的分歧也可能消失。我们提出了实验性模型类修订,其中发现策略联合提出结构编辑和诊断实验,以测试该编辑是否必要。该方法将类级可区分性目标(其中共享参数化必须解释所有选定的实验)与随时有效的序贯证据相结合,仅在当前类被拒绝后才触发结构修订。在400个保留的受控动态环境中,联合策略在32个真实实验的预算下达到89.5%的精确恢复率,比最强匹配基线提高了10.0个百分点,同时需要更少的执行实验和候选拟合。学习到的修订-实验配对可跨未见过的机制组合、保留但可表达的基元、参数外推和转移的实验成本进行迁移;当真实机制不在编辑语法内时,它以5.5%的虚假支持率在88%的情况下检测到库不足。修订增益也转移到ODEBench和ODEBase模型库任务以及DiscoverPhysics世界中。这些结果支持一种科学发现观点,即决定理论应使哪些机制可表达以及在哪里收集证据被视为一个单一的序贯决策问题。
英文摘要:
Scientific discovery systems typically optimize experiments within a fixed hypothesis space. This creates a failure mode when all available candidates omit the same missing mechanism: candidate disagreement can collapse even while the model class is systematically wrong. We formulate experimental model-class revision, in which a discovery policy jointly proposes a structural edit and a diagnostic experiment that tests whether that edit is necessary. The method couples a class-level distinguishability objective, in which one shared parameterization must explain all selected experiments, with anytime-valid sequential evidence that triggers structural revision only after the current class is rejected. On 400 held-out controlled dynamical environments, the joint policy reaches 89.5% exact recovery with a budget of 32 real experiments, improving the strongest matched baseline by 10.0 percentage points while requiring fewer executed experiments and candidate fits. The learned revision-experiment pairing transfers across unseen mechanism combinations, held-out but expressible primitives, parameter extrapolation, and shifted experiment costs; when the true mechanism is outside the edit grammar, it detects library insufficiency in 88% of cases with a 5.5% false-support rate. Revision gains also transfer to ODEBench and ODEBase model-library tasks, as well as DiscoverPhysics worlds. These results support a view of scientific discovery in which deciding what mechanisms a theory should make expressible and where to collect evidence are treated as a single sequential decision problem.