arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.08475cs.ROcs.CVcs.LG

DexMan:从人类和生成视频中学习双臂灵巧操作

DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated Videos

Jhen Hsieh, Kuan-Hsun Tu, Kuo-Han Hung, Tsung-Wei Ke

首次发表 更新
浏览论文内容

中文总结 AI 辅助

DexMan是一个自动化框架,从人类和生成视频中学习双臂灵巧操作,无需标定或标注,通过接触奖励提升策略,在姿态估计和成功率上达到最先进水平。

中文摘要 AI 辅助

我们提出了DexMan,一个自动化框架,将人类视觉演示转化为仿真中的人形机器人双臂灵巧操作技能。DexMan直接处理人类操作刚体物体的第三人称视频,无需相机标定、深度传感器、扫描的3D物体资产或真实的手和物体运动标注。与先前仅考虑简化浮动手的方方法不同,它直接控制人形机器人,并利用基于接触的新型奖励,从野外视频中估计的嘈杂手-物体姿态中改进策略学习。DexMan在TACO基准上的物体姿态估计中达到了最先进的性能,在ADD-S和VSD上分别取得了0.08和0.12的绝对提升。同时,其强化学习策略在OakInk-v2上的成功率比先前方法高出19%。此外,DexMan可以从真实和合成视频中生成技能,无需手动数据收集和昂贵的动作捕捉,从而能够创建大规模、多样化的数据集,用于训练通用灵巧操作。

英文摘要

We present DexMan, an automated framework that converts human visual demonstrations into bimanual dexterous manipulation skills for humanoid robots in simulation. Operating directly on third-person videos of humans manipulating rigid objects, DexMan eliminates the need for camera calibration, depth sensors, scanned 3D object assets, or ground-truth hand and object motion annotations. Unlike prior approaches that consider only simplified floating hands, it directly controls a humanoid robot and leverages novel contact-based rewards to improve policy learning from noisy hand-object poses estimated from in-the-wild videos. DexMan achieves state-of-the-art performance in object pose estimation on the TACO benchmark, with absolute gains of 0.08 and 0.12 in ADD-S and VSD. Meanwhile, its reinforcement learning policy surpasses previous methods by 19% in success rate on OakInk-v2. Furthermore, DexMan can generate skills from both real and synthetic videos, without the need for manual data collection and costly motion capture, and enabling the creation of large-scale, diverse datasets for training generalist dexterous manipulation.

发表机构

  • National Taiwan University(国立台湾大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑