发表机构
ACE Robotics; CUHK; CPII under InnoHK; Tongji University; Shanghai Jiao Tong University; Zhejiang University; Nanyang Technological University(ACE机器人公司; 香港中文大学; InnoHK框架下的CPII机构; 同济大学; 上海交通大学; 浙江大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对杂乱场景下灵巧抓取数据匮乏问题,提出种子-过滤策略构建含260万场景的基准,引入OmniDex模型结合Soft Winner-Takes-All学习与物理约束,实现最优性能与强泛化性。
AI 中文摘要
灵巧抓取是具身智能的基础原语,需要海量数据来训练鲁棒模型。由于真实世界数据采集成本高昂,仿真已成为主流范式。然而,尽管杂乱场景最能反映真实应用,但在其中学习抓取却面临大规模数据严重匮乏的瓶颈。为解决这一问题,我们整理了高质量的3D物体和支撑底座,提出了一种可扩展的种子-过滤策略,该策略绕过了缓慢的场景级优化。这产生了一个前所未有的基准,包含超过260万个场景和0.4B个场景特定的抓取真值,具有多样的真实布局,搭配丰富的语义和几何观测。此外,我们引入了OmniDex模型,以克服当前生成模型所面临的抓取多模态和最后一毫米精度误差问题。通过在训练过程中将Soft Winner-Takes-All学习与受人类启发的物理约束相结合,并利用物理驱动的排序,我们的方法实现了鲁棒的灵巧抓取,且没有后优化的延迟。实验结果表明,OmniDex模型在多样场景、视角和未见物体上实现了最先进的性能和强大的泛化能力。
英文摘要
Dexterous grasping is the foundational primitive in embodied AI, demanding massive data to train robust models. As real-world data collection is expensive, simulation has become the mainstream paradigm. Yet, while cluttered scenes best reflect real-world applications, learning to grasp within them is bottlenecked by a critical scarcity of large-scale data. To resolve this, we curate high-quality 3D objects and supporting bases, proposing a scalable seed-and-filter strategy that bypasses sluggish scene-level optimization. This yields an unprecedented benchmark comprising over 2.6 million scenes and 0.4B scene-specific grasp ground truths, featuring diverse realistic layouts paired with rich semantic and geometric observations. Furthermore, we introduce the OmniDex model to overcome the grasp multimodality and last-millimeter precision errors plaguing current generative models. By coupling Soft Winner-Takes-All learning with human-inspired physical constraints during training, and utilizing physics-driven ranking, our approach achieves robust dexterous grasping without the latency of post-optimization. Experimental results show that OmniDex model achieves state-of-the-art performance and strong generalization across diverse scenes, views, and unseen objects.