CADGrasp: 学习接触和碰撞感知的通用灵巧抓取在杂乱场景中
CADGrasp: Learning Contact and Collision Aware General Dexterous Grasping in Cluttered Scenes
- Center on Frontiers of Computing Studies, School of Computer Science, Peking University(前沿计算研究中心,计算机学院,北京大学)
- National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家级重点实验室,计算机学院,北京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
CADGrasp通过两阶段算法实现接触和碰撞感知的灵巧抓取,利用点云输入预测稀疏IBS表示,并结合占用扩散模型和能量函数优化,提升在杂乱场景中的抓取成功率与稳定性。
AI中文摘要:
在杂乱环境中进行灵巧抓取面临巨大挑战,这源于灵巧手的高自由度、遮挡以及由于不同物体几何形状和复杂布局可能产生的碰撞。为了解决这些挑战,我们提出了CADGrasp,一种使用单视角点云输入进行通用灵巧抓取的两阶段算法。在第一阶段,我们预测稀疏IBS,一种场景解耦、接触和碰撞感知的表示,作为优化目标。稀疏IBS紧凑地编码了灵巧手与场景之间的几何和接触关系,使得能够实现稳定且无碰撞的灵巧抓取姿态优化。为了增强对这种高维表示的预测,我们引入了带有体素级条件引导和力闭合分数过滤的占用扩散模型。在第二阶段,我们开发了几种能量函数和排序策略,基于稀疏IBS进行优化,以生成高质量的灵巧抓取姿态。在模拟和现实世界中的广泛实验验证了我们方法的有效性,证明了其在减少碰撞的同时,能够保持在各种物体和复杂场景中较高的抓取成功率。
英文摘要:
Dexterous grasping in cluttered environments presents substantial challenges due to the high degrees of freedom of dexterous hands, occlusion, and potential collisions arising from diverse object geometries and complex layouts. To address these challenges, we propose CADGrasp, a two-stage algorithm for general dexterous grasping using single-view point cloud inputs. In the first stage, we predict sparse IBS, a scene-decoupled, contact- and collision-aware representation, as the optimization target. Sparse IBS compactly encodes the geometric and contact relationships between the dexterous hand and the scene, enabling stable and collision-free dexterous grasp pose optimization. To enhance the prediction of this high-dimensional representation, we introduce an occupancy-diffusion model with voxel-level conditional guidance and force closure score filtering. In the second stage, we develop several energy functions and ranking strategies for optimization based on sparse IBS to generate high-quality dexterous grasp poses. Extensive experiments in both simulated and real-world settings validate the effectiveness of our approach, demonstrating its capability to mitigate collisions while maintaining a high grasp success rate across diverse objects and complex scenes.