发表机构
Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过重复开放式自主材料发现实验,发现不同策略虽导致筛选和构建结构数量差异,但最终收敛于相同材料前沿,同时揭示共享输入导致的常见模式错误。
AI 中文摘要
科学智能体大多根据其是否完成任务或恢复已知结果来评估;我们转而研究重复开放式活动中出现的差异。十六个独立初始化的会话,使用相同的模型-框架配置,接收了包含12,499种金属有机框架的冻结数据库、一个甲烷储存目标、固定协议和一周的预算。策略分化为四种方法,筛选了100至5,000种结构,其中八个会话构建了2,253种假设结构。然而,这些智能体恢复的材料前沿接近200 cm^3/cm^3,并且对数据库多孔区域的独立计算发现,其九个最佳结构均出现在它们的报告中。对一半智能体实施的强制检查将新运行的重现率从八分之一提高到八分之八,但未能显著提高结论的有效性,因为十六个智能体中有十五个选择了相同的审计排除条目,即一个不完整的结构,其缺失的阴离子造成了虚假的孔隙体积。因此,重复的智能体既揭示了稳健的结论,也揭示了共享输入导致的常见模式错误。
英文摘要
Scientific agents are mostly evaluated on whether they complete tasks or recover known results; we instead study variation across repeated open-ended campaigns. Sixteen separately initialized sessions of one model-harness configuration received a frozen database of 12,499 metal-organic frameworks, a methane-storage objective, a pinned protocol and a one-week budget. Strategies diverged into four approaches spanning 100--5,000 screened structures, and eight built 2,253 hypothetical structures. Yet the agents recovered the same materials frontier near 200 cm^3/cm^3, and an independent calculation of the database's porous region found its nine best structures all among their reports. Enforced checks on half the agents raised fresh-run reproduction from one of eight to eight of eight but could not detectably improve conclusion validity, because fifteen of sixteen agents selected the same audit-excluded entry, an incomplete structure whose missing anions created artificial pore volume. Replicated agents thus reveal both robust conclusions and common-mode errors from shared inputs.