arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18948cs.RO

RoboEdit:将人类操作视频转化为可扩展的机器人体验

RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience

Yaowei Guo, Zeng Tao, Yuxin Jiang, Yunuo Chen, Zhiyang Dou, Yuxiang Ma, Yin Yang, Demetri Terzopoulos, Ying Jiang, Chenfanfu Jiang

AI总结:

研究人员开发RoboEdit工具套件,通过RoboEdit-ADC自动流程构建含17.4万对齐视频的数据集,利用RoboEdit-Trans引擎实现人类操作视频到机器人视频的高质量转化,为机器人学习提供可扩展监督。

AI中文摘要:

收集机器人手物交互数据成本高昂且与机器人本体特异性相关,然而大量存在的人类物体视频却无法用于机器人训练。本文提出RoboEdit,这是一套将人类操作视频转化为动作一致、物理合理且带有对齐3D手部状态的机器人视频的视频编辑工具套件。为实现可扩展的监督,我们引入RoboEdit-ADC,这是一种跨本体从RGB视频重建并重定向3D交互的自动流程。该流程生成RoboEdit-14M,这是一个包含17.4万对对齐视频(共1400万帧)的大规模数据集,涵盖7种机器人本体、多样化场景及交互类型。核心编辑引擎RoboEdit-Trans采用跨本体适配模块,在适配外观与运动的同时保持时间一致性,还集成了3D机器人状态解码器以恢复每帧手部状态,用于结构化运动监督。实验表明,RoboEdit实现了最先进的编辑质量,并支持下游机器人控制策略完成真实世界操作任务。最终,RoboEdit套件释放了未标注人类视频的巨大潜力,为可泛化的机器人学习提供可扩展、高保真的视觉与3D运动监督。

英文摘要:

Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning. Project webpage: https://roboedit.github.io/

补充信息

↑