SparseDFF:用于一次性灵巧操作的稀疏视图特征蒸馏
SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
- CFCS, School of Computer Science, Peking University(北京大学计算机学院CFCS)
- Institute for AI, Peking University(北京大学人工智能研究院)
- Department of Computer Science, Stanford University(斯坦福大学计算机科学系)
- PKU-WUHAN Institute for Artificial Intelligence, China(北京大学武汉人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出SparseDFF,利用大型2D视觉模型从稀疏RGBD图像提取语义特征生成3D稠密特征场,通过特征细化网络和对比损失实现一次性灵巧操作,在真实场景中验证了对刚性和可变形物体的泛化能力。
AI中文摘要:
人类在将操作能力迁移到不同形状、姿态和外观的物体上时表现出非凡的技巧,这种能力源于他们对不同实例之间语义对应关系的理解。为了使机器人具备类似的高级理解能力,我们提出了SparseDFF,一种新颖的3D场景DFF(稠密特征场),利用大型2D视觉模型从稀疏RGBD图像中提取语义特征,尽管该领域与许多固定相机设置的任务相关,但研究仍然有限。SparseDFF生成视图一致的3D DFF,通过将图像特征映射到3D点云,实现一次性学习灵巧操作。SparseDFF的核心是一个特征细化网络,通过视图间的对比损失和点剪枝机制优化特征连续性。这有助于最小化相对于末端执行器参数的特征差异,桥接演示与目标操作。在真实世界场景中使用灵巧手进行验证,SparseDFF在操作刚性和可变形物体方面均表现出有效性,展示了跨物体和场景变化的显著泛化能力。
英文摘要:
Humans demonstrate remarkable skill in transferring manipulation abilities across objects of varying shapes, poses, and appearances, a capability rooted in their understanding of semantic correspondences between different instances. To equip robots with a similar high-level comprehension, we present SparseDFF, a novel DFF for 3D scenes utilizing large 2D vision models to extract semantic features from sparse RGBD images, a domain where research is limited despite its relevance to many tasks with fixed-camera setups. SparseDFF generates view-consistent 3D DFFs, enabling efficient one-shot learning of dexterous manipulations by mapping image features to a 3D point cloud. Central to SparseDFF is a feature refinement network, optimized with a contrastive loss between views and a point-pruning mechanism for feature continuity. This facilitates the minimization of feature discrepancies w.r.t. end-effector parameters, bridging demonstrations and target manipulations. Validated in real-world scenarios with a dexterous hand, SparseDFF proves effective in manipulating both rigid and deformable objects, demonstrating significant generalization capabilities across object and scene variations.