arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BioVLN:生物医学实验室视觉语言导航的仿真平台

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

Zhe Liu, Quan Lu, Zhaohui Du, Zhe Wang, Huanbo Jin, Jiaming Gu, Qi Wang, Ting Xiao, Minting Pan, Dongzhan Zhou

arXiv 2607.26914首次发表:更新:

发表机构

East China University of Science and Technology; Ruijin Hospital, Shanghai Jiao Tong University School of Medicine; Shanghai AI Laboratory(华东理工大学; 上海交通大学医学院附属瑞金医院; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有导航平台不适用于实验室仪器的问题,提出BioVLN仿真平台,通过分区域建模仪器实现更精准导航,实验验证其提升成功率并降低安全风险。

AI 中文摘要

生物医学实验室机器人必须在执行实验程序前导航至仪器处。现有具身导航平台针对家庭环境设计,将目标视为物体中心或任意附近位置,该表示不适用于实验室仪器——需从操作侧接近,同时与周围设备保持安全间距。我们提出BioVLN,用于开发和评估生物医学实验室视觉语言导航智能体的仿真平台。BioVLN将每个仪器分为三个区域:物理主体、周围安全区域、可用侧前方的操作区域,该模型统一应用于场景生成、目标放置、导航评估和安全分析,因此成功取决于到达可访问仪器的位置。BioVLN支持程序式场景生成和手动设计环境,生成47个场景和1667个 episode。标准化导航和强化学习接口支持轨迹收集和策略训练。实验显示,几何探索的成功率达74.4%–87.5%,而在操作区域采样多个有效位置可将成功率提升至83.3%–92.5%并减少不安全接近。

英文摘要

Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inadequate for laboratory instruments, which must be approached from their operating side while maintaining safe clearance from surrounding equipment. We introduce BioVLN, a simulation platform for developing and evaluating visual-language navigation agents in biomedical laboratories. BioVLN represents each instrument with three regions: its physical body, a surrounding clearance region, and an operation area in front of the usable side. This model is applied consistently to scene generation, target placement, navigation evaluation, and safety analysis, so success depends on reaching a position from which the instrument can be accessed. BioVLN supports procedural scene generation and manually designed environments, producing 47 scenes and 1667 episodes. Standardized navigation and reinforcement-learning interfaces enable trajectory collection and policy training. Experiments show that geometric exploration reaches 74.4--87.5% success, while sampling multiple valid positions in the operation area improves success to 83.3--92.5% and reduces unsafe proximity.

Comments17 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑