V2-STRep:基于VLM的结构化任务表示,从生成视频中获取可复用机器人技能
V2-STRep: VLM-Grounded Structured Task Representations for Reusable Robot Skills Acquired from Generated Videos
浏览论文内容
中文总结 AI 辅助
V2-STRep通过VLM生成结构化任务表示,将生成视频中的运动转化为可复用机器人技能,实现零样本跨场景迁移,并在六个真实任务中提升执行成功率。
中文摘要 AI 辅助
人类操作视频为获取机器人技能提供了丰富的运动和交互线索,无需机器人示范。视频生成模型从初始场景图像和任务指令中合成此类示范,避免了为每个任务录制示范的需求。然而,恢复的运动仅捕获一种场景特定的实现,任务结构、几何关系和约束仍然隐含。我们提出V2-STRep,一种零样本框架,通过基于VLM的结构化任务表示,将生成的视频运动转换为可复用的机器人技能。该表示指定运动阶段、参考和任务相关约束,目标由最小几何结构描述:点、点法线、轴、平面和完整6D位姿。VLM提供的2D图像空间线索利用RGB-D观测提升至3D,以重建任务几何和候选抓取位姿。几何特定规则将运动迁移到新场景,而任务约束的轨迹优化将抓取选择与完整机器人运动规划耦合。它保留任务要求,同时利用剩余旋转自由度以适应关节限制。更新部署基础和约束可在不生成另一个视频的情况下,在新的兼容指令下实现复用。在六个真实世界操作任务上的实验表明,与基线相比执行成功率提高,成功获取的技能可靠跨场景迁移,并适应更改的部署指令。
英文摘要
Human manipulation videos provide rich motion and interaction cues for acquiring robot skills without robot demonstrations. Video generation models synthesize such demonstrations from an initial scene image and task instruction, avoiding the need to record demonstrations for each task. However, the recovered motion captures only one scene-specific realization, leaving task structure, geometric relations, and constraints implicit. We present V2-STRep, a zero-shot framework that converts generated video motion into reusable robot skills through VLM-grounded structured task representations. The representation specifies motion phases, references, and task-relevant constraints, with targets described by minimal geometric structures: points, point-normals, axes, planes, and full 6D poses. VLM-provided 2D image-space cues are lifted into 3D using RGB-D observations to reconstruct task geometry and candidate grasp poses. Geometry-specific rules transfer motion to new scenes, while task-constrained trajectory optimization couples grasp selection with complete robot motion planning. It preserves task requirements while using remaining rotational freedom to accommodate joint limits. Updating deployment grounding and constraints enables reuse under new compatible instructions without generating another video. Experiments on six real-world manipulation tasks demonstrate improved execution success over baselines, reliable cross-scene transfer of successfully acquired skills, and adaptation to changed deployment instructions.
发表机构
- Technische Universität Wien (TU Wien)(维也纳工业大学)
- German Aerospace Center (DLR)(德国航空航天中心)
机构由 AI 辅助整理,请以论文原文为准。