发表机构
School of Computer Science & Informatics, University of Liverpool; Department of Chemistry, University of Liverpool(利物浦大学计算机科学与信息学院; 利物浦大学化学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对机器人化学实验的安全风险,提出SAFE-CHEM不确定性感知框架,通过混合控制架构在不确定性超阈值时切换至备份控制器,提升任务成功率并减少安全违规,实现零样本仿真到真实的迁移。
AI 中文摘要
自主机器人系统在化学实验室中的部署正在加快实验流程,并为AI驱动的科学发现提供基础数据。然而,尽管数据驱动方法在掌握灵巧技能方面取得了成功,安全性仍是其在高风险领域(如早期材料化学实验)部署的主要障碍。具体而言,基于学习的策略常常难以区分安全与不安全的动作,导致过度自信的外推,可能引发灾难性故障。为缓解这些安全风险,我们提出SAFE-CHEM,一种专为稳健的基于学习的机器人化学家设计的不确定性感知框架。我们的方法利用一组基于循环神经网络的模仿学习策略集成,通过动作预测的方差在线量化认知不确定性。通过使用核密度估计表征该方差的成功条件密度,我们引入一种混合控制架构,当不确定性超过校准的安全阈值时,自动从学习策略切换到确定性的基于规则的备份控制器。我们在三个基础实验室操作任务中评估SAFE-CHEM,实证结果表明,与传统的单策略基线相比,该混合策略提高了整体任务成功率并减少了关键安全违规。最后,我们通过零样本仿真到真实迁移,在物理Franka Production 3机械臂上验证了该框架的实际可行性。
英文摘要
The deployment of autonomous robotic systems in chemistry laboratories is accelerating experimental workflows and providing the foundational data for AI-driven scientific discovery. However, despite the success of data-driven methods in acquiring dexterous skills, safety remains a primary barrier to their deployment in high-risk domains, such as early-stage materials chemistry experiments. Specifically, learning-based policies frequently struggle to distinguish between safe and unsafe actions, leading to overconfident extrapolation and potentially catastrophic failures. To mitigate these safety risks, we propose SAFE-CHEM, an uncertainty-aware framework designed for robust, learning-based robotic chemists. Our approach leverages an ensemble of recurrent neural network-based imitation learning policies to quantify epistemic uncertainty online through the variance of action predictions. By characterising the success-conditioned density of this variance using kernel density estimation, we introduce a hybrid control architecture that autonomously switches from the learned policy to a deterministic, rule-based backup controller when uncertainty exceeds a calibrated safety threshold. We evaluate SAFE-CHEM across three fundamental laboratory manipulation tasks, where our empirical results demonstrate that this hybrid strategy improves overall task success rates and reduces critical safety violations compared to traditional single-policy baselines. Finally, we demonstrate the practical viability of the framework through zero-shot sim-to-real transfer onto a physical Franka Production 3 robot manipulator.