arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23704cs.ROcs.CV

LabRobFail:化学自动驾驶实验室中机器人故障分析的基准

LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory

Haobo Wang, Baoli Sun, Anqi Zou, Dongsheng Huang, Zelin Lv, Ning Wang, Rui Li, Dongzhan Zhou, Weiyu Guo, Zhihui Wang, Wanli Ouyang

AI总结:

针对化学自动驾驶实验室中机器人可靠性受化学实验特性限制、故障数据稀缺及评估协议缺乏等问题,引入LabRobFail框架,通过注入故障构建数据集,开发专门模型,提升故障检测等能力及下游任务成功率。

AI中文摘要:

在自动驾驶实验室中部署实体智能体可加速科学发现,但化学实验的不可逆性和安全关键性质限制了其可靠性。由于故障数据稀缺且缺乏细粒度评估协议,进展受阻。为此引入LabRobFail框架,LabRobFail-Sim在控制、物理和语义层面注入可控故障以构建包含超20000条轨迹的LabRobFail-Data。LabRobFail-Bench评估六种能力。还开发了LabRobFail-VLM,在可见环境中表现出色,集成后提高了下游VLA任务成功率,证明了细粒度故障理解对闭环恢复和可靠实验室自主性的价值。

英文摘要:

The deployment of embodied agents in self-driving laboratories could accelerate scientific discovery, yet their reliability is constrained by the irreversible and safety-critical nature of chemical experiments. Progress is further hindered by scarce failure data and the lack of fine-grained evaluation protocols. To address these challenges, we introduce LabRobFail, a failure-centric framework for learning and evaluating robotic failure analysis in chemical laboratories. LabRobFail-Sim injects controllable failures at the control, physics, and semantic levels, enabling the construction of LabRobFail-Data, which contains over 20,000 trajectories across 70+ task scenarios, five failure categories, and 11 fine-grained failure types. LabRobFail-Bench evaluates six capabilities spanning task understanding, failure detection, temporal localization, severity assessment, failure classification, and actionable correction. We further develop LabRobFail-VLM, a domain-specialized vision-language model that generates structured failure diagnoses and recovery instructions. On seen environments, it achieves 90.83% failure-detection accuracy and 77.21% temporal-localization accuracy, substantially outperforming general-purpose VLMs. When integrated as a real-time supervisor, it improves downstream task success rates by 4-16 percentage points, demonstrating the value of fine-grained failure understanding for closed-loop recovery and reliable laboratory autonomy. Our code and data are available at https://github.com/Su-ISE-2001/SciRobo

补充信息

↑