arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniHOI:从单目人类视频实现灵巧手-物体交互

OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video

Ting Mao, Yanming Shao, Ziheng Wang, Haoyu Liu, Yiqun Wang, Xuanye Wu, Yao Mu

arXiv 2610.10855首次发表:更新:

发表机构

Zhejiang University; The University of Hong Kong; Shanghai Jiao Tong University; Shanghai AI Laboratory(浙江大学; 香港大学; 上海交通大学; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OmniHOI是一种将手-物体交互RGB视频转换为灵巧手轨迹的管道,通过多阶段物理一致性优化,在动作捕捉和单目视频迁移任务中均优于现有方法,可在真实双臂机器人上执行。

AI 中文摘要

人类操作的单目视频提供了丰富的灵巧操作演示,但从单视角重建手-物体交互并将其迁移至机器人手仍存在困难,限制了其直接用于机器人执行。现有方法要么需要特定任务的强化学习(RL)训练,可扩展性受限;要么假设使用干净的动作捕捉轨迹,因此无法直接在视频上运行。我们提出OmniHOI,这一管道可将手-物体交互的RGB视频转换为与交互一致的灵巧手轨迹。核心思路是利用各阶段可用的证据强制执行物理一致性:重建阶段使用图像证据,重定向阶段使用接触几何,物理在环优化阶段使用动力学。每个阶段直接优化对应的表示,在错误向下游传播或被学习策略吸收前进行修正。在150条动作捕捉轨迹被迁移至5种自由度(DoF)为6至22的灵巧手的实验中,我们的方法成功率达39%-89%,而现有迁移方法的最高成功率仅为31%;在60个单目视频片段上,我们的方法成功率为53%,而现有最优的视频到机器人管道的成功率为28%。其生成的轨迹还可在真实的双臂机器人上执行,完成各类任务。

英文摘要

Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution. Prior methods either require task-specific RL training, limiting scalability, or assume clean motion-capture trajectories and thus cannot operate directly on video. We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands. The key idea is to enforce physical consistency using the evidence available at each stage: image evidence during reconstruction, contact geometry during retargeting, and dynamics during physics-in-the-loop refinement. Each stage optimizes the corresponding representation directly, correcting errors before they propagate downstream or must be absorbed by a learned policy. Across 150 motion-capture trajectories transferred to each of five dexterous hands with 6 to 22 DoF, we achieve 39-89% success, compared with at most 31% for prior transfer methods. On 60 monocular video clips, we achieve 53% success, compared with 28% for the best prior video-to-robot pipeline. Its trajectories also execute on a real bimanual robot across diverse tasks.

CommentsThe project page is available at https://hydrofoooil.github.io/OmniHOI/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑