arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09365cs.RO

PhysV2A:用于视频到机器人操作的可达性门控和语义掩码约束的可行性完成

PhysV2A: Reachability-Gated and Semantic-Mask-Constrained Feasibility Completion for Video-to-Robot Manipulation

Haohui Huang, Junda Duan, Tao Teng, Chenguang Yang

首次发表
浏览论文内容

中文总结 AI 辅助

研究如何将视频衍生的6D对象运动转换为机器人可执行轨迹,提出PhysV2A框架,通过可达性门控选择和语义掩码约束细化可操作性,经真实机器人实验验证,该框架能提升任务成功率、减少运动可行性失败并产生更好轨迹。

中文摘要 AI 辅助

基于视频的操作可从人类演示、生成的视频或RGB-D观察中提供以对象为中心的运动先验,但这些先验通常与具体实现无关,无法由特定机器人直接执行。本文提出了PhysV2A,这是一个用于将视频衍生的6D对象运动转换为机器人可执行操作轨迹的可达性门控和语义掩码约束的可行性完成框架。关键思想是将抓取可行性视为轨迹条件而非局部条件。PhysV2A进行分层可达性门控选择,对于选定的可达轨迹,通过VLM辅助和规则验证的S-Mask识别任务关键和可松弛的笛卡尔组件,通过冗余优先优化和有界笛卡尔松弛实现语义掩码约束的可操作性细化。在四个桌面操作任务上的真实机器人实验表明,PhysV2A提高了任务成功率,减少了运动可行性失败,并产生了具有有界语义偏差的更好条件的轨迹。

英文摘要

Video-based manipulation provides object-centric motion priors from human demonstrations, generated videos, or RGB-D observations, but such priors are typically embodiment-agnostic and cannot be directly executed by a specific robot. This paper presents \textbf{PhysV2A}, a reachability-gated and semantic-mask-constrained feasibility-completion framework for converting video-derived 6D object motion into robot-executable manipulation trajectories. The key idea is to treat grasp feasibility as trajectory-conditioned rather than local: each RGB-D-generated 6-DoF grasp candidate is rigidly coupled with the recovered object motion to form a grasp-conditioned TCP trajectory hypothesis. PhysV2A then performs hierarchical reachability-gated selection, where infeasible grasp--trajectory pairs are rejected by robot-centric kinematic checks and surviving candidates are ranked by downstream execution suitability. For the selected reachable trajectory, a VLM-assisted and rule-validated S-Mask identifies task-critical and relaxable Cartesian components, enabling semantic-mask-constrained manipulability refinement through redundancy-first optimization and bounded Cartesian relaxation. Real-robot experiments on four tabletop manipulation tasks show that PhysV2A improves task success over representative video-prior and IK-only baselines, reduces kinematic-feasibility failures, and produces better-conditioned trajectories with bounded semantic deviations.

发表机构

  • School of Automation, Guangdong University of Technology(广东工业大学自动化学院)
  • University of Liverpool(利物浦大学)
  • Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算学系)

机构由 AI 辅助整理,请以论文原文为准。

↑