arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.09491cs.ROcs.CV

DCIRNet:用于透明和反射物体灵巧抓取的带迭代细化的深度补全

DCIRNet: Depth Completion with Iterative Refinement for Dexterous Grasping of Transparent and Reflective Objects

  • State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Guanghu Xie, Zhiduo Jiang, Yonglong Zhang, Yang Liu, Zongwu Xie, Baoshi Cao, Hong Liu

更新

AI总结:

提出多模态深度补全网络 DCIRNet,通过 RGB-深度特征融合及多阶段监督迭代细化恢复透明、反射物体深度,并将其用于灵巧抓取,成功率提升44%。

AI中文摘要:

日常环境中的透明和反射物体由于其独特的视觉特性(如镜面反射和光透射),给深度传感器带来了重大挑战。这些特性通常会导致不完整或不准确的深度估计,从而严重影响下游基于几何的视觉任务,包括物体识别、场景重建和机器人操作。为了解决透明和反射物体中深度信息缺失的问题,我们提出了 DCIRNet,这是一种新颖的多模态深度补全网络,可有效整合 RGB 图像和深度图以提升深度估计质量。我们的方法包含一个创新的多模态特征融合模块,旨在提取 RGB 图像与不完整深度图之间的互补信息。此外,我们引入了多阶段监督和深度细化策略,可逐步改进深度补全,并有效缓解物体边界模糊的问题。我们将深度补全模型集成到灵巧抓取框架中,使透明和反射物体的抓取成功率提高了 44%。我们在公共数据集上进行了大量实验,DCIRNet 在其中展现出卓越性能。实验结果验证了我们方法的有效性,并证实了其在各类透明和反射物体上的强大泛化能力。

英文摘要:

Transparent and reflective objects in everyday environments pose significant challenges for depth sensors due to their unique visual properties, such as specular reflections and light transmission. These characteristics often lead to incomplete or inaccurate depth estimation, which severely impacts downstream geometry-based vision tasks, including object recognition, scene reconstruction, and robotic manipulation. To address the issue of missing depth information in transparent and reflective objects, we propose DCIRNet, a novel multimodal depth completion network that effectively integrates RGB images and depth maps to enhance depth estimation quality. Our approach incorporates an innovative multimodal feature fusion module designed to extract complementary information between RGB images and incomplete depth maps. Furthermore, we introduce a multi-stage supervision and depth refinement strategy that progressively improves depth completion and effectively mitigates the issue of blurred object boundaries. We integrate our depth completion model into dexterous grasping frameworks and achieve a $44\%$ improvement in the grasp success rate for transparent and reflective objects. We conduct extensive experiments on public datasets, where DCIRNet demonstrates superior performance. The experimental results validate the effectiveness of our approach and confirm its strong generalization capability across various transparent and reflective objects.

↑