发表机构
Texas A&M University; Polara Labs Inc.(得克萨斯农工大学; Polara实验室公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出可识别性驱动实验智能体LLM-IDEA,配备可识别性引擎判定平台期类型,在多个基准测试中验证其能提升自主发现机制世界模型的效率与准确性。
AI 中文摘要
大型语言模型智能体正越来越多地被部署为自主科学家,在极少人类监督下设计实验并推断机制世界模型。然而可识别性常被忽视:当达到平台期时,智能体需要知道是自身能力不足,还是模型根本无法从数据中识别,此时无论进行多少同类额外实验都无济于事。我们提出了可识别性驱动实验智能体(LLM-IDEA),用于闭环发现,其配备的可识别性引擎会返回三类平台期判定:能力极限、在设计类别内可解决,或已被证实穷尽。在ODEBench上,62个含自由常数的系统中有60个在第0轮即可识别;RC电路在协议可运行的所有实验中均被证实穷尽,而收获模型可通过添加一个初始条件解决。该可识别性引擎在Lotka-Volterra、Van der Pol、Lorenz及药代动力学模型上重现了已知判定,其中它推荐了药理学家使用的静脉注射分支,且在四种观测设计中将我们基准的深度评分者排在最后。在DiscoverPhysics基准上,它发现了两个公共世界,其解释规则奖励了合法实验无法实现的区分,且所有轨迹准确的模型均未通过解释等级(15个模型全部未通过,而可识别世界中9个模型有5个通过,p=0.012)。在我们提出的双体测试台“外星宇宙”中,力定律在可证非可识别与可识别协议间切换,LLM-IDEA在可识别协议上的8个随机种子中有8个达到至少3的发现深度,而无该引擎时仅1个。因此,自主发现智能体可计算而非猜测平台期是需要更多搜索、同类更好实验,还是不同类型实验。
英文摘要
Large language model agents are being increasingly deployed as autonomous scientists, designing experiments and inferring mechanistic world models with minimal human oversight. Yet identifiability is often overlooked: when a plateau is reached, the agent needs to know whether it is not yet capable enough or the model simply is not identifiable from the data, in which case no amount of further experimentation of the same kind can help. We propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted. On ODEBench, 60 of the 62 systems with free constants are identifiable at round 0; the RC circuit is certified exhausted for every experiment that protocol can run, and a harvesting model is resolvable by one added initial condition. The identifiability engine reproduces known verdicts on Lotka-Volterra, Van der Pol, Lorenz, and a pharmacokinetic model, where it recommends the intravenous arm pharmacologists use, and it ranks the depth scorer of our own benchmark last among four observation designs. On the DiscoverPhysics benchmark, it finds two public worlds whose explanation rubric rewards a distinction no legal experiment can make, and every model there with accurate trajectories failed the explanation grade (15 of 15, against 5 of 9 in identifiable worlds, p = 0.012). On the Alien Universe, a two-body testbed we propose in which a force law switches between a provably non-identifiable and an identifiable protocol, LLM-IDEA on the identifiable protocol reaches discovery depth at least three on 8/8 seeds versus 1/8 without it. An autonomous discovery agent can thus compute, rather than guess, whether a plateau calls for more search, a better experiment of the same kind, or a different kind of experiment.