arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TransBiolab:一个杂乱透明生物医学物体的真实世界多视图数据集

TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects

Ke Ma, Yifei Wang, Meng Wang, Tian Xia

arXiv 2607.21071首次发表:更新:

发表机构

School of Artificial Intelligence and Automation, Huazhong University of Science and Technology; College of Design and Innovation, Tongji University; Shanghai Institute for Intelligent Autonomous Systems, Tongji University; School of Software and Engineering, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院; 同济大学设计创意学院; 同济大学上海智能无人系统研究院; 华中科技大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自主生物医学实验室透明物体视觉感知数据稀缺问题,提出TrainsBiolab真实世界多视图数据集,含多物体杂乱、遮挡场景数据及多种注释,定义相关基准并评估,为自主实验室操作视觉任务提供资源。

AI 中文摘要

自主生物医学实验室越来越依赖视觉感知来识别、定位和操作透明塑料制品,但针对此场景的高质量真实世界数据集仍然有限。在杂乱的多物体场景中,领域相关数据的稀缺尤其具有限制性,即使对于当代视觉基础模型,相互遮挡和视图相关的外观变化仍然具有挑战性。现有的透明物体数据集有先进的分割、深度和姿态估计,但通常不评估多物体杂乱、遮挡和校准多视图捕获的组合设置。为填补这一空白,我们展示了TrainsBiolab,一个作为校准多视图序列捕获的杂乱透明生物医学物体的真实世界RGB-D数据集。它包含来自98个场景的161315帧和超过15种实验室物体类型的103万个实例注释,包括6D姿态、完整和可见掩码、深度以及每帧相机校准。该数据集沿反映操作难度的三个轴组织:物体类别、一帧中物体的总数和相机视点。我们进一步定义了以数据集为中心的分割、深度估计和完成以及6D姿态估计基准,并报告了由发布的注释和校准实现的系统级机器人操作评估。通过关注重复的透明实例、杂乱和多视图实验室捕获,TrainsBiolab为自主实验室操作中的分割、深度估计、6D姿态估计和多视图推理提供了资源。

英文摘要

Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quality real-world datasets for this setting remain limited. The scarcity of domain-relevant data is particularly restrictive in cluttered multi-object scenes, where mutual occlusion and view-dependent appearance changes remain challenging even for contemporary visual foundation models. Existing transparent-object datasets have advanced segmentation, depth, and pose estimation, but they usually do not evaluate the combined setting of multi-object clutter, occlusion, and calibrated multi-view capture that characterizes real laboratory manipulation scenes. To address this gap, we present TrainsBiolab, a real-world RGB-D dataset of cluttered transparent biomedical objects captured as calibrated multi-view sequences. TrainsBiolab contains 161,315 frames from 98 scenes and 1.03M instance annotations over 15 laboratory object types, including 6D poses, full and visible masks, depth, and per-frame camera calibration. The dataset is organized along three axes that reflect operational difficulty: object category, the total number of objects in a frame, and camera viewpoint. We further define dataset-centric benchmarks for segmentation, depth estimation and completion, and 6D pose estimation, and report a system-level robot manipulation evaluation enabled by the released annotations and calibrations. By focusing on repeated transparent instances, clutter, and multi-view laboratory capture, TrainsBiolab provides a resource for segmentation, depth estimation, 6D pose estimation, and multi-view reasoning in autonomous laboratory manipulation. Project page: https://dualtransparency.github.io/TransBiolab/.

Comments9 pages, 10 figures, accepted by ACM Multimedia 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑