发表机构
University of Massachusetts Amherst; Genesis AI(马萨诸塞大学阿默斯特分校; Genesis AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
IM-ENGINE利用图像编辑作为中间表示,结合仿真器根基,为机器人学习生成兼具语义意义和物理可执行性的操控监督数据,应用于灵巧抓取与目标状态生成。
AI 中文摘要
基于学习的操控需要既具有语义意义又可在物理上执行的监督信号,但当前的数据流水线通常仅能提供其中一种属性。人类演示能够捕捉意图,但采集成本高昂且受限于人机具身差异,而仿真虽可扩展数据生成,却往往对功能性行为描述不足。我们提出IM-ENGINE,一种以仿真器为根基的流水线,将图像编辑用作具身数据生成的中间表示。给定具有已知几何、深度、分割和相机参数的渲染场景,IM-ENGINE通过编辑图像注入任务相关语义,利用仿真器先验和未改变的锚定物体恢复显式3D状态,在物理引擎中细化该状态,并将其转换为机器人可执行的监督信号。我们针对灵巧抓取综合和目标状态生成实例化了该流水线。对于抓取,IM-ENGINE在图像空间中生成人类抓取,恢复手-物体交互,将其重定向到机器人手,并细化为经过物理验证的机器人抓取。对于目标生成,它将渲染场景编辑为期望结果,恢复目标物体姿态,并细化为物理上有效、语义上有意义的目标和轨迹。这种生成式语义先验与仿真器根基的结合,为机器人学习实现了可扩展的任务相关监督。
英文摘要
Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can scale data generation but often under-specifies functional behavior. We present IM-ENGINE, a simulator-grounded pipeline that uses image editing as an intermediate representation for embodied data generation. Given a rendered scene with known geometry, depth, segmentation, and camera parameters, IM-ENGINE edits the image to inject task-relevant semantics, recovers explicit 3D state using simulator priors and an unchanged anchor object, refines the state in physics, and converts it into robot-executable supervision. We instantiate the pipeline for dexterous grasp synthesis and goal-state generation. For grasping, IM-ENGINE generates a human grasp in image space, recovers the hand-object interaction, retargets it to a robot hand, and refines it into physically validated robot grasps. For goal generation, it edits a rendered scene into a desired outcome, recovers the target-object pose, and refines it into physically valid, semantically meaningful goals and trajectories. This combination of generative semantic priors and simulator grounding enables scalable task-relevant supervision for robot learning.
Comments26 pages, 19 figures, 3 tables