arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AnyViewDex: 基于RGB观测的视角不变灵巧操作

AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations

Soham Patil, Om Sanjay Gunjal, Sourabh Bhosale, Arhan Chavare, Ramandeep Singh Hora, Spandan Roy

arXiv 2609.20107首次发表:更新:

发表机构

Robotics Research Center, IIIT-H; Veermata Jijabai Technological Institute Mumbai; Indian Institute of Science Education and Research Bhopal(IIIT-H机器人研究中心; 孟买维尔马塔·吉贾拜技术学院; 印度科学教育与研究学院博帕尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出AnyViewDex,通过仿真中结合多视角对比对齐与3D几何监督,实现仅用单目RGB零样本的视角不变灵巧操作,硬件实验验证了其有效性。

AI 中文摘要

多指灵巧操作的视觉运动策略对相机视角变化高度敏感。为实现视角不变性,近期方法日益依赖显式3D模态,如RGB-D或点云,这可能在真实世界部署中引入硬件依赖、校准要求以及对传感器噪声的脆弱性。在本工作中,我们展示了通过在仿真期间将几何知识编码到视觉表示中,可以在无需显式测试时3D感知的情况下实现视角不变控制。我们提出AnyViewDex,一种非对称训练流程,结合了多视角对比对齐与特权3D几何监督。通过在仿真训练期间回归绝对3D物体坐标,该辅助目标提供了几何接地信号,缓解了全局池化对比嵌入的空间坍缩。在部署时,策略仅使用未校准的单目RGB和本体感觉即可零样本运行。我们在强化学习和师生蒸馏中均验证了该方法。在配备16自由度LEAP手的xArm7上的硬件评估中,AnyViewDex在八个未见物体和六个未校准视角下达到76.7%的抓取成功率(480次试验;所有消融条件下共2400次),表明几何接地的单目策略在无需测试时深度的情况下零样本迁移。项目页面:此https URL

英文摘要

Visuomotor policies for multi-fingered dexterous manipulation are highly sensitive to camera viewpoint shifts. To achieve view invariance, recent methods increasingly rely on explicit 3D modalities like RGB-D or point clouds, which can introduce hardware dependencies, calibration requirements, and vulnerability to sensor noise during real-world deployment. In this work, we show that view-invariant control can be achieved without explicit test-time 3D sensing by encoding geometric knowledge into the visual representation during simulation. We present AnyViewDex, an asymmetric training pipeline that combines multi-view contrastive alignment with privileged 3D geometric supervision. By regressing absolute 3D object coordinates during simulated training, this auxiliary objective provides a geometric grounding signal that mitigates the spatial collapse of the globally pooled contrastive embedding. At deployment, the policy operates zero-shot using only uncalibrated monocular RGB and proprioception. We validate this approach across both reinforcement learning and student-teacher distillation. In hardware evaluation on an xArm7 with a 16-DoF LEAP Hand, AnyViewDex reaches 76.7% grasping success across eight unseen objects and six uncalibrated viewpoints (480 trials; 2,400 across all ablation conditions), indicating that geometrically grounded monocular policies transfer zero-shot without test-time depth. Project Page: https://anyviewdex.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑