发表机构
Stanford University; Nvidia; Microsoft Research(斯坦福大学; 英伟达; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RoboRender提出面向机器人的视频生成框架,利用仿真深度、语言和掩码生成逼真RGB视频,训练策略实现零样本现实部署,成功率大幅提升。
AI 中文摘要
仿真能够实现大规模、低成本的机器人数据生成,但由于仿真到现实的视觉差异,在仿真中训练的策略往往难以迁移到现实世界。现有方法通常依赖中间表示,这可能会丢弃丰富的语义信息或在部署时需要额外的感知模块。我们通过RoboRender解决这一视觉仿真到现实差距,RoboRender是一个将仿真轨迹转换为逼真RGB视频以用于策略学习的框架。RoboRender训练一个面向机器人的视频生成模型,该模型以仿真深度视频、语言指令和机器人RGB掩码视频为条件,在保留仿真器几何、机器人运动和动作标签的同时,合成逼真的纹理、背景和干扰物。生成的RGB视频与仿真器提供的状态和动作配对,用于训练策略以实现零样本现实部署。在机器人视频测试集上,我们的视频模型在生成质量上优于基于深度条件的视频生成基线。在涉及拾取放置、铰接物体操作和移动操作任务的现实实验中,基于RoboRender生成数据训练的策略实现了71%的平均成功率,分别比原始仿真渲染和传统视觉域随机化高出约7.1倍和3.6倍。我们进一步表明,每条仿真轨迹生成的视频越多,策略性能越好,使打开任务的成功率提高了65个百分点。这些结果表明,生成式视频渲染缓解了零样本策略迁移的视觉仿真到现实差距。项目网站:此https URL。
英文摘要
Simulation enables large-scale, low-cost robot data generation, but policies trained in simulation often fail to transfer to the real world due to the sim-to-real visual discrepancies. Existing approaches often rely on intermediate representations, which can discard rich semantic information or require additional perception modules at deployment. We address this visual sim-to-real gap with RoboRender, a framework that converts simulated trajectories into photorealistic RGB videos for policy learning. RoboRender trains a robot-oriented video generation model conditioned on simulated depth videos, language instructions, and robot RGB mask videos, preserving simulator geometry, robot motion, and action labels while synthesizing realistic textures, backgrounds, and distractors. The generated RGB videos are paired with simulator-provided states and actions to train policies for zero-shot real-world deployment. On robot video test sets, our video model outperforms depth-conditioned video generation baselines in generation quality. In real-world experiments across pick-and-place, articulated-object manipulation, and mobile manipulation tasks, policies trained on RoboRender-generated data achieve a 71% average success rate, outperforming raw simulation renderings and conventional visual domain randomization by approximately 7.1x and 3.6x, respectively. We further show that policy performance improves with more generated videos per simulation trajectory, increasing opening-task success by 65 percentage points. These results demonstrate that generative video rendering mitigates the visual sim-to-real gap for zero-shot policy transfer. Project website: https://robo-render.github.io/.