arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OccluDex:自遮挡下第一人称灵巧操作的分层三维视觉-触觉表示学习

OccluDex: Hierarchical 3D Visuo-Tactile Representation Learning for Egocentric Dexterous Manipulation under Self-Occlusion

Ziheng Xu, Yueyuan Chen, Xinyuan He, Guoxing Liu, Yuanshuo Tan, Huiming Pan, Bin He, Shuo Jiang, Peter B. Shull

arXiv 2609.39017首次发表:更新:

发表机构

Shanghai Jiao Tong University; Tongji University; The Chinese University of Hong Kong, Shenzhen(上海交通大学; 同济大学; 香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对第一人称灵巧操作中手部自遮挡导致视觉信息不足的问题,提出分层视觉-触觉表示学习框架OccluDex,通过多尺度掩码自编码和跨模态注意力融合,在仿真和物理实验中显著提升操作准确率并实现零样本泛化。

AI 中文摘要

可靠的灵巧操作需要在交互过程中持续估计物体几何形状和手-物体接触状态。然而,在第一人称感知下,操作手经常遮挡任务相关的物体表面和接触区域,减少了可用于状态估计的视觉证据,从而使鲁棒闭环控制和对未见物体几何形状的泛化变得尤为困难。为解决这一问题,我们提出了OccluDex,一种分层三维视觉-触觉表示学习框架,该框架整合全局几何结构与局部接触信息,以在手-物体交互过程中的动态自遮挡下实现鲁棒操作。OccluDex采用多尺度掩码自编码逐步编码部分三维几何,并通过跨模态注意力将触觉接触令牌与高层几何特征融合。编码器从同步的人类视觉-触觉演示中预训练,并作为冻结的感知骨干迁移到下游强化学习。我们在水龙头旋转任务(要求完整顺时针旋转手柄一圈)和桌面物体重新定向任务(要求180度重新定向且不倾倒)上评估了OccluDex。在仿真实验中,OccluDex在未见物体上的准确率比最强的最先进基线模型高12.6%,在已见物体上高8.3%。此外,我们使用Shadow Hand进行了物理实验,展示了在未见物理物体上的成功零样本仿真到现实泛化。这些结果能够使人形机器人即使在操作手遮挡视觉的情况下,也能对已见和未见物体进行第一人称物体操作。

英文摘要

Reliable dexterous manipulation requires continuous estimation of object geometry and hand-object contact throughout interaction. With egocentric sensing, however, the manipulating hand frequently occludes task-relevant object surfaces and contact regions, reducing the visual evidence available for state estimation and thereby making robust closed-loop control and generalization to unseen object geometries particularly challenging. To address this, we present OccluDex, a hierarchical 3D visuo-tactile representation learning framework that integrates global geometric structure with local contact information for robust manipulation under dynamic self-occlusion during hand-object interaction. OccluDex adopts multi-scale masked autoencoding to progressively encode partial 3D geometry and fuses tactile contact tokens with high-level geometric features through cross-modal attention. The encoder is pretrained from synchronized human visuo-tactile demonstrations and transferred as a frozen perceptual backbone for downstream reinforcement learning. We evaluate OccluDex on a faucet rotation task, requiring one full clockwise handle revolution, and a tabletop object reorientation task, requiring a 180-degree tabletop object reorientation without toppling. In simulation experiments, OccluDex demonstrated 12.6% higher accuracy for unseen objects and 8.3% higher accuracy for previously seen objects than the strongest state-of-the-art baseline models. Physical experiments were further performed with a Shadow Hand to demonstrate successful zero-shot sim-to-real generalization on unseen physical objects. This results could enable humanoid egocentric object manipulation for seen and unseen objects even when the manipulating robotic hand occludes vision.

Comments8 pages, 5 figures. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑