arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将人类视频机器人化:实现物理一致性的交互

Robotizing Human Videos with Physically Consistent Interactions

Ching-Lam Cheng, Shengfeng He, Bin Zhu

arXiv 2610.06137首次发表:更新:

发表机构

Singapore Management University(新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种物理一致性交互方法,通过接触重建与深度感知合成将人类视频机器人化,与机器人演示共同训练,在RoboTwin任务中取得最高平均成功率,并增强对视觉干扰的鲁棒性。

AI 中文摘要

人类视频提供了可扩展的操作数据,但人类手部与机器人操作器之间的具身差距限制了其直接使用。现有的视频编辑方法用渲染的机器人替换手部,然而不准确的交互重建和合成可能导致不一致的抓取和不合逻辑的机器人-物体遮挡。我们从两个互补的物理方面解决这些失败:交互几何和场景可见性。首先,一个交互感知的接触重建模块结合手-物体分割与网格级接触预测,以恢复密集的3D接触,然后将其转换为时间稳定的平行颚夹爪抓取。其次,一个深度感知的合成模块利用场景和机器人深度来强制实现物理一致的机器人-物体遮挡。生成的视频以机器人兼容的形式保留了人类演示的交互结构,并与机器人演示共同训练。使用相同的人类视频和机器人数据,我们与仅机器人训练和原始Masquerade流程进行比较。在四个RoboTwin任务和两个Diffusion Policy视觉编码器中,我们的方法实现了最高的平均成功率,尤其在分布外场景变化下表现突出。实际部署进一步表明,当任务几何可观测时,所提出的共同训练方法提高了对视觉干扰物的鲁棒性,而深度敏感抓取的性能仍受限于单摄像头设置。

英文摘要

Human videos offer scalable manipulation data, but the embodiment gap between human hands and robot manipulators limits their direct use. Existing video-editing methods replace hands with rendered robots, yet inaccurate interaction reconstruction and compositing can produce inconsistent grasps and implausible robot-object occlusions. We address these failures from two complementary physical aspects: interaction geometry and scene visibility. First, an interaction-aware contact reconstruction module combines hand-object segmentation with mesh-level contact prediction to recover dense 3D contacts, then converts them into temporally stabilized grasps for parallel-jaw grippers. Second, a depth-aware compositing module uses scene and robot depth to enforce physically consistent robot-object occlusions. The resulting videos preserve the interaction structure of human demonstrations in a robot-compatible form and are co-trained with robot demonstrations. Using identical human videos and robot data, we compare against robot-only training and the original Masquerade pipeline. Across four RoboTwin tasks and two Diffusion Policy visual encoders, our method achieves the highest average success rates, with especially strong gains under out-of-distribution scene variation. Real-world deployment further shows that the proposed co-training approach improves robustness to visual distractors when the task geometry is observable, while performance on depth-sensitive grasps remains limited by the single-camera setup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑