arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在机器人化学实验室中对大型语言模型智能体进行压力测试

Stress-testing large language model agents in a robotic chemistry laboratory

Lulu Guo, Yingkai Sun, Xiaobo Li, Luyao Ge, Ziming Wang, Haitao Zheng, Jingyu Li, Huijuan Zhang, Bingxu Chen, Daobin Liu, Yuebo Liu, Jie Li, Xiaohui Li, Linjiang Chen, Yi Luo, Jun Jiang

arXiv 2607.23045首次发表:更新:

AI 中文总结

研究以机器人化学实验室为测试平台,使科学智能可衡量,通过4608次试验发现工作流程执行率低、长期规划有挑战,反馈促使局部调整,提供了部署评估及改进诊断框架。

AI 中文摘要

人工智能通过知识、推理和计划生成来评估,然而科学智能需要可靠的物理行动和对证据的适应能力。在此,我们使用机器人化学实验室作为物理世界测试平台,以使科学智能可衡量。其45个模块化工作站作为机器可读技能进行展示,共进行了4608次试验。在实验室限制下,只有3.3%的试验产生了专家评估的可执行工作流程;即使是最佳系统也仅达到28.1%。长期规划仍是挑战:只有三个可执行工作流程超过30个操作,最长的包含44个。经过五轮实验,反馈促使进行局部调整,但没有工作流程级的重新规划或分析方法的重新设计。通过使物理可执行性和基于证据的重新规划可衡量,我们的研究为部署准备情况提供了基于证据的评估,并为指导朝着物理基础的自主研究进行闭环改进提供了诊断框架。

英文摘要

AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency measurable. Its 45 modular workstations exposed as machine-readable skills enabled 4,608 trials. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints; even the best system achieved 28.1%. Long-horizon planning remained a challenge: only three executable workflows exceeded 30 operations, although the longest contained 44. Across five rounds, experimental feedback prompted local adjustments but no workflow-level replanning or analytical-method redesign. By making physical executability and evidence-driven replanning measurable, our study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements towards physically grounded autonomous research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑