发表机构
Robotics Research Center, IIIT Hyderabad; IIT-Jodhpur(海得拉巴国际信息技术学院机器人研究中心; 印度理工学院焦特布尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GlassFormer,融合毫米波雷达与RGB-D实现实时玻璃分割,利用雷达反射生成空间先验,在低光下较纯视觉方法显著提升鲁棒性。
AI 中文摘要
透明表面在建筑环境中无处不在,但它们仍然是机器人感知中持续存在的失败案例。RGB相机感知到的是玻璃后面的背景而非表面本身,而深度传感器如LiDAR、飞行时间(ToF)和RGB-D在透明区域往往返回无效或背景测量值。因此,仅依赖光学传感的系统可能将玻璃墙、门或镜子误判为自由空间,从而危及安全可靠的导航。现有的玻璃分割方法通过从RGB图像中学习反射、边界和语义上下文等视觉线索来解决这一问题。尽管在有利的照明和视角条件下有效,但这些线索在低光环境、强光下或玻璃表面无特征或部分遮挡时会退化。在这项工作中,我们提出了一种多模态框架,将毫米波雷达与RGB-D传感融合,用于实时透明表面分割。雷达在玻璃表面反射强烈,提供了一种几何线索,在视觉和深度失效的地方恰好保持可靠。我们利用这种跨模态不一致性生成雷达引导的空间先验,并通过跨模态注意力将其集成到轻量级基于Transformer的分割网络GlassFormer中。我们在覆盖所有场景类型的混合条件测试分割和专门设计用于压力测试纯视觉方法的低光分割上报告了结果。GlassFormer在混合分割上达到0.88 mIoU,在低光分割上达到0.59 mIoU,展示了相对于纯视觉基线的显著鲁棒性提升,同时在资源受限平台上保持实时性能。
英文摘要
Transparent surfaces are ubiquitous in built environments, yet they remain a persistent failure case for robotic perception. RGB cameras perceive the background behind glass rather than the surface itself, while depth sensors such as LiDAR, time-of-flight, and RGB-D often return invalid or background measurements in transparent regions. As a result, systems that rely solely on optical sensing may misinterpret glass walls, doors, or mirrors as free space, compromising safe and reliable navigation. Existing glass segmentation approaches address this by learning visual cues such as reflections, boundaries, and semantic context from RGB images. While effective under favourable lighting and viewing conditions, these cues degrade in low-light environments, under glare, or when glass surfaces are featureless or partially occluded. In this work, we propose a multimodal framework that fuses millimetre-wave radar with RGB-D sensing for real-time transparent surface segmentation. Radar reflects strongly off glass surfaces, providing a geometric cue that remains reliable precisely where vision and depth fail. We exploit this cross-modal inconsistency to generate a radar-guided spatial prior, which is integrated into a lightweight transformer-based segmentation network, GlassFormer, via cross-modal attention. We report results on a mixed-condition test split covering all scene types and a dedicated low-light split designed to stress vision-only methods. GlassFormer achieves 0.88 mIoU on the mixed split, and 0.59 mIoU on the low light split, demonstrating substantial robustness gains over vision-only baselines while maintaining real-time performance on resource-constrained platforms.
CommentsAccepted for presentation at IEEE IROS 2026. Code available at https://github.com/Suhani92/GlassFormer