arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21229cs.ROcs.AI

KnowDemo:从人类视频生成知识引导的机器人演示

KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos

Zhiyuan Gao, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Schäfer, Michael Beetz

首次发表
浏览论文内容

中文总结 AI 辅助

KnowDemo利用人类视频中的结构化操作知识,通过VLM提取与推理生成多样化机器人演示,提升候选规划成功率,并在仿真数据上微调π0.5模型实现三任务的仿真到现实迁移。

中文摘要 AI 辅助

学习机器人操作策略通常需要大量的演示数据,而这些数据在真实机器人上采集成本高昂。近期的方法通过适配恢复的运动并在仿真中验证所得轨迹,从人类视频生成机器人演示。然而,以运动参考适配为中心的方法可能因保留演示中的接触策略和子任务顺序而限制行为多样性,同时,对任务需求和场景关系理解不足会因生成无效候选而降低演示生成效率。为解决这些局限,我们提出KnowDemo,一个利用来自人类视频的结构化操作知识为目标工作区生成多样化机器人演示的框架。为区分任务需求与演示特定选择,我们开发了一个基于视觉语言模型(VLM)的知识提取与推理模块,该模块将对象和动作描述与推断的任务条件、演示参考及允许的执行变化相关联。为将这些知识转化为可执行的演示,我们针对目标场景实体和几何解析描述,以在运动规划和仿真之前指导候选生成与筛选。所得演示通过替代接触策略和有效的子任务顺序展现多模态行为,并带有结构化执行标签。实验表明,与仅参考配置相比,增加了验证过的执行模式,并通过任务引导的抓取采样提高了候选规划成功率。为验证生成数据用于策略学习,我们在仿真数据上微调预训练的$\pi_{0.5}$模型,在三个任务上实现了仿真到现实的迁移。项目页面:此https URL

英文摘要

Learning robot manipulation policies typically requires substantial demonstration data, which are costly to collect on real robots. Recent methods generate robot demonstrations from human videos by adapting recovered motion and validating the resulting trajectories in simulation. However, methods centered on motion-reference adaptation can limit behavioral diversity by retaining the demonstrated contact strategies and subtask orders, while insufficient understanding of task requirements and scene relations can reduce demonstration generation efficiency by generating invalid candidates. To address these limitations, we propose KnowDemo, a framework that uses structured manipulation knowledge from human videos to generate diverse robot demonstrations for a target workspace. To distinguish task requirements from demonstration-specific choices, we develop a knowledge extraction and reasoning module based on a vision-language model (VLM) that associates object and action descriptions with inferred task conditions, demonstration references, and permissible execution variations. To translate this knowledge into executable demonstrations, we resolve the descriptions against target-scene entities and geometry to guide candidate generation and screening before motion planning and simulation. The resulting demonstrations exhibit multimodal behavior through alternative contact strategies and valid subtask orders, with structured execution labels. Experiments demonstrate additional verified execution modes beyond a reference-only configuration and improved candidate planning success through task-guided grasp sampling. To validate the generated data for policy learning, we fine-tune the pretrained $π_{0.5}$ model on simulation data, achieving sim-to-real transfer across three tasks. Project page: https://zhiyuan-gao.github.io/knowdemo/

发表机构

  • University of Bremen(不来梅大学)
  • Robotics Institute Germany (RIG)(德国机器人研究所(RIG))

机构由 AI 辅助整理,请以论文原文为准。

↑