发表机构
School of Automation, Guangdong University of Technology; University of Liverpool; Department of Computing, The Hong Kong Polytechnic University; State Key Laboratory of Submarine Geoscience, School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(广东工业大学自动化学院; 利物浦大学; 香港理工大学计算学系; 上海交通大学海洋地球科学国家重点实验室,自动化与智能感知学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在将生成视频转换为机器人可执行操作轨迹。核心方法是GenVid2Robot框架,通过采样语义锚点、跟踪验证运动等步骤实现。主要贡献是提高了生成视频引导操作的可靠性,通过多种手段建立视觉运动先验。
AI 中文摘要
生成的视频为机器人操作提供了有用的视觉运动先验,但视觉上的合理性并不意味着物理上的可执行性。生成的视频通常缺乏度量几何、抓取基础、机器人运动学可行性和执行时反馈,这使得直接轨迹重放在实际操作中不可靠。本文提出了GenVid2Robot,这是一个刚性几何一致性框架,可将生成的视频运动转换为可执行的真实机器人操作轨迹。给定初始RGB-D观察和任务指令,GenVid2Robot从真实第一帧中采样与任务相关的语义锚点,通过生成的视频候选跟踪这些锚点,并验证生成的2D运动是否可以由稀疏相对$SE(3)$模型下的第一帧RGB-D锚点解释。只有几何上一致的运动才会传输到机器人。然后将接受的相对运动应用于通过掩码约束抓取选择的真实抓取时TCP姿态,生成与视觉运动先验和物理抓取配置一致的抓取条件执行轨迹。为了减少由RGB-D噪声、校准残差和小接触引起的位移导致的执行不匹配,有界深度补偿模块在不假设完全在线重新规划的情况下校正局部深度方向误差。真实机器人实验表明,GenVid2Robot通过用稀疏度量几何、抓取约束、机器人可行性检查和有界执行反馈来建立视觉运动先验,提高了生成视频引导操作的可靠性。
英文摘要
Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot kinematic feasibility, and execution-time feedback, which makes direct trajectory replay unreliable in real-world manipulation. This paper presents GenVid2Robot, a rigid-geometric consistency framework that converts generated video motion into executable real-robot manipulation trajectories. Given an initial RGB-D observation and a task instruction, GenVid2Robot samples task-relevant semantic anchors from the real first frame, tracks these anchors through generated video candidates, and verifies whether the resulting 2D motion can be explained by first-frame RGB-D anchors under a sparse relative $SE(3)$ model. In this way, generated videos are treated as uncertain visual motion hypotheses rather than direct robot demonstrations. Only geometrically consistent motion is transferred to the robot. The accepted relative motion is then applied to the real grasp-time TCP pose selected by mask-constrained grasping, producing a grasp-conditioned execution trajectory that is consistent with both the visual motion prior and the physical grasp configuration. To reduce execution mismatch caused by RGB-D noise, calibration residuals, and small contact-induced displacement, a bounded depth-compensation module corrects local depth-direction errors without assuming full online replanning. Real-robot experiments demonstrate that GenVid2Robot improves the reliability of generated-video-guided manipulation by grounding visual motion priors with sparse metric geometry, grasp constraints, robot feasibility checking, and bounded execution feedback.
CommentsPreprint