GRASPLAT:通过新视角合成实现灵巧抓取
GRASPLAT: Enabling dexterous grasping through novel view synthesis
- TeV, Fondazione Bruno Kessler(TeV,布鲁诺·凯瑟尔基金会)
- PAVIS, Fondazione Istituto Italiano di Tecnologia(PAVIS,意大利技术研究院基金会)
- ISR, Instituto Superior Técnico, Universidade de Lisboa(ISR,里斯本大学理工学院)
- University of Trento(特伦托大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
GRASPLAT提出仅用RGB图像训练、结合3D高斯泼溅合成新视角并引入光度损失,提升多指手灵巧抓取成功率,最高提升36.9%。
AI中文摘要:
实现多指手的灵巧机器人抓取仍然是一个重大挑战。虽然现有方法依赖完整的3D扫描来预测抓取姿态,但这些方法因在现实场景中难以获取高质量3D数据而受到限制。在本文中,我们提出了GRASPLAT,一种新颖的抓取框架,它利用一致的3D信息,但仅使用RGB图像进行训练。我们的关键见解是,通过合成手抓取物体的物理上合理的图像,我们可以回归出成功抓取对应的手部关节。为实现这一目标,我们利用3D高斯泼溅(3D Gaussian Splatting)生成真实手-物交互的高保真新视角,从而实现基于RGB数据的端到端训练。与先前方法不同,我们的方法引入了光度损失(photometric loss),通过最小化渲染图像与真实图像之间的差异来细化抓取预测。我们在合成和真实世界的抓取数据集上进行了大量实验,证明GRASPLAT在现有基于图像的方法基础上将抓取成功率提高了高达36.9%。项目页面:https://mbortolon97.github.io/grasplat/
英文摘要:
Achieving dexterous robotic grasping with multi-fingered hands remains a significant challenge. While existing methods rely on complete 3D scans to predict grasp poses, these approaches face limitations due to the difficulty of acquiring high-quality 3D data in real-world scenarios. In this paper, we introduce GRASPLAT, a novel grasping framework that leverages consistent 3D information while being trained solely on RGB images. Our key insight is that by synthesizing physically plausible images of a hand grasping an object, we can regress the corresponding hand joints for a successful grasp. To achieve this, we utilize 3D Gaussian Splatting to generate high-fidelity novel views of real hand-object interactions, enabling end-to-end training with RGB data. Unlike prior methods, our approach incorporates a photometric loss that refines grasp predictions by minimizing discrepancies between rendered and real images. We conduct extensive experiments on both synthetic and real-world grasping datasets, demonstrating that GRASPLAT improves grasp success rates up to 36.9% over existing image-based methods. Project page: https://mbortolon97.github.io/grasplat/