arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19119cs.CV

追踪、表达、行动:从日常人体视频生成关节运动

Track, Articulate, Act: Generating Articulation from Casual Human Videos

Jiaming Zhang, Homanga Bharadhwaj

AI总结:

本文提出一个从单目RGB视频重建关节物体及其手-物交互的真实到模拟框架,利用密集3D点轨迹推断关节,并复用预训练模型实现分割、重建与交互重放。

AI中文摘要:

人体视频包含丰富的机器人操作因果证据:它们揭示了手部运动如何引起物体运动,并产生物体状态中与任务相关的改变。在这项工作中,我们研究诸如门、抽屉、橱柜、笔记本电脑、烤箱和铰链容器等关节物体,这些物体在日常生活中无处不在,并为具身交互带来了独特的挑战。这些物体不能由单一姿态表示;它们的运动取决于底层部件和关节。我们引入了一个真实到模拟的框架,该框架从随意的单目RGB视频中重建一个可用于模拟的关节物体和手-物交互,无需RGB-D或多视角输入、先验扫描、手动指定的关节或机器人演示。我们的关键见解是,密集的3D点轨迹提供了一种与具身无关的关节线索:固定连杆上的点近似保持静止,而移动连杆上的点则遵循一致的旋转或棱柱运动。我们的方法分割连杆,估计关节及其状态轨迹,重建关节资产,并将恢复的3D手部运动与物体对齐。我们方法的核心是一个模块化配方,它重新利用强大的预训练模型进行单图像3D重建、网格分割和3D场景流,通过显式几何推理连接它们的预测以推断关节。我们使用重建的关节物体和人类手部轨迹,通过MuJoCo中的接触来重放交互。该框架展示了预训练视觉模型和显式运动推理如何将随意的人体视频转化为适用于下游具身交互的关节物体模型。此https URL

英文摘要:

Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes in object state. In this work, we study articulated objects such as doors, drawers, cabinets, laptops, ovens, and hinged containers that are ubiquitous in daily life and present unique challenges for embodied interaction. These objects cannot be represented by a single pose; their motion depends on the underlying parts and joints. We introduce a real-to-sim framework that reconstructs a simulation-ready articulated object and hand-object interaction from a casual monocular RGB video, without RGB-D or multi-view input, prior scans, manually specified joints, or robot demonstrations. Our key insight is that dense 3D point tracks provide an embodiment-agnostic articulation cue: points on the fixed link remain approximately stationary, while points on the moving link follow coherent revolute or prismatic motion. Our method segments the links, estimates the joint and its state trajectory, reconstructs an articulated asset, and aligns the recovered 3D hand motion with the object. Central to our approach is a modular recipe that repurposes powerful pretrained models for single-image 3D reconstruction, mesh segmentation, and 3D scene flow, connecting their predictions through explicit geometric reasoning to infer articulation. We use the reconstructed articulated object and the human hand trajectory to replay interactions through contact in MuJoCo. The framework shows how pretrained vision models and explicit motion reasoning can turn casual human videos into articulated object models suitable for downstream embodied interactions. https://track-articulate-act.github.io/

补充信息

↑