arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探索,然后提交:利用语言模型进行测量高效的科学定律发现

Explore, Then Commit: Measurement-Efficient Scientific Law Discovery with Language Models

Kautik Mandve, Dileepa Fernando

arXiv 2610.07620首次发表:更新:

发表机构

Vernon Hills High School; Singapore University of Technology and Design(弗农希尔斯高中; 新加坡科技设计大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出探索-然后提交协议,利用语言模型假设、程序化规划器测量和综合提示,在NewtonBench上实现测量高效的科学定律发现,显著减少测量次数并降低误差。

AI 中文摘要

科学定律发现需要选择测量并将证据转化为控制方程。我们评估了一种探索-然后提交协议,其中大型语言模型提出假设,程序化规划器收集测量,而新的提示从固定观测中综合最终定律。该协议结合了结构化探测、自动数值诊断、受限测量批次和可选解释器访问。在576次NewtonBench试验中,我们使用GPT-4.1-mini和中等难度的GPT-4.1复现,在12个物理模块上比较了八种配置。在中等任务上,启用解释器的规划器在GPT-4.1-mini上每次试验使用8.6次测量对比22.5次,在GPT-4.1上使用8.9次对比43.0次。它们的平均基于幅度的均方根对数误差分别从2.514降至0.202,从0.626降至0.149。额外的审计在覆盖敏感分析中保留了不完整和无效的提交。观察到的符号准确性增益在各模块间不太一致,随机采集与分歧评分具有竞争力。测量节省发生在每个模块中,但不平等的批次约束阻止将其仅归因于采集质量。这些结果支持完整协议作为一种有前途的测量高效配置,同时将其因果组件和超越无噪声直接方程任务的泛化问题留待解决。

英文摘要

Scientific law discovery requires selecting measurements and converting evidence into a governing equation. We evaluate an explore-then-commit protocol in which a large language model proposes hypotheses, a programmatic planner gathers measurements, and a fresh prompt synthesizes the final law from fixed observations. The protocol combines structured probes, automatic numerical diagnostics, restricted measurement batches, and optional interpreter access. Across 576 NewtonBench trials, we compare eight configurations on 12 physics modules using GPT-4.1-mini and a medium-difficulty GPT-4.1 replication. On medium tasks, interpreter-enabled planners use 8.6 versus 22.5 measurements per trial for GPT-4.1-mini and 8.9 versus 43.0 for GPT-4.1. Their mean magnitude-based root-mean-squared logarithmic error falls from 2.514 to 0.202 and from 0.626 to 0.149, respectively. An additional audit retains incomplete and invalid submissions in a coverage-sensitive analysis. Observed symbolic-accuracy gains are less consistent across modules, and random acquisition is competitive with disagreement scoring. Measurement savings occur in every module, but unequal batch constraints prevent attributing them solely to acquisition quality. These results support the complete protocol as a promising measurement-efficient configuration, while leaving its causal components and generalization beyond noiseless direct-equation tasks unresolved.

Comments19 pages, including Supplementary Material S1; code and data included as ancillary files. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑