发表机构
Technical University of Munich; Peking University; Huawei Heisenberg Research Center; Huawei CloudRobo Lab; Nanjing University(慕尼黑工业大学; 北京大学; 华为海森堡研究中心; 华为云机器人实验室; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VidAct提出从单目视频学习以对象为中心的3D感知操作策略,通过对象网格重建、残差轨迹迁移和完整点云辅助任务,实现零样本真实部署。
AI 中文摘要
视频演示为学习操作提供了一种可扩展的替代方案,以替代昂贵的机器人数据,然而现有的基于重建的方法通常依赖于受限的相机视角或人机重定向,同时重建的轨迹难以适应新的物体配置而不扭曲轨迹形状。另一个关键限制是,由此产生的策略往往缺乏精确的对象级3D几何感知,限制了对象锚定和对象形状感知,而这些对于精确操作至关重要。为弥合这些差距,我们提出了VidAct,一个高效的视频到机器人框架,它从每个任务的单目视频中学习以对象为中心的、具有3D意识的操作策略,并实现零样本的真实世界部署。VidAct由三个关键组件组成。首先,VidAct从任意演示视频中重建对象网格和运动,并在静态对象坐标系中对运动进行规范化,避免了特定于实体的重定向,并适应了多样的相机视角。其次,VidAct采用残差轨迹迁移来将重建的运动适应到新的对象配置,同时保持其运动形状。最后,作为关键的策略学习组件,VidAct在每一帧预测模拟提供的特权完整对象点云作为辅助任务,同时保持仅使用RGB的部署,从而在对象姿态和3D几何上提供密集的对象中心监督。在人类、机器人、生成和互联网视频上的实验证明了广泛的视频适用性和零样本部署。逐帧的完整对象3D监督提高了策略泛化和模拟到现实的成功率,而残差轨迹迁移实现了可靠的轨迹适应,并具有更好的形状保持。
英文摘要
Video demonstrations offer a scalable alternative to costly robot data for learning manipulation, yet existing reconstruction-based approaches often rely on constrained camera viewpoints or human-to-robot retargeting, while the reconstructed trajectories are difficult to adapt to new objects configurations without distorting the trajectory shape. Another key limitation is that the resulting policies often lack precise object-level 3D geometry awareness, limiting object grounding and object shape awareness critical for precise manipulation. To bridge these gaps, we propose VidAct, an efficient video-to-robot framework that learns object-centric, 3D-aware manipulation policies from a single monocular video per task and enables zero-shot real-world deployment. VidAct consists of three key components. First, VidAct reconstructs object meshes and motion from arbitrary demo videos and canonicalizes the motion in the static object frame, avoiding embodiment-specific retargeting and accommodating diverse camera viewpoints. Second, VidAct employ residual trajectory transfer for adapting the reconstructed motion to novel object configurations while preserving its motion shape. Finally, as the key policy-learning component, VidAct predicts simulation-provided privileged complete-object point clouds at each frame as an auxiliary task while retaining RGB-only deployment, providing dense object-centric supervision over both object pose and 3D geometry. Experiments on human, robot, generated, and internet videos demonstrate broad video applicability and zero-shot deployment. Per-frame complete-object 3D supervision improves policy generalization and sim-to-real success, while residual trajectory transfer enables reliable trajectory adaptation with better shape preservation.