arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19661cs.RO

ReShoot:用于视觉运动策略学习的机器人录制的演示的生成式视觉域随机化

ReShoot: Generative Visual Domain Randomization of Recorded Robot Demonstrations for Visuomotor Policy Learning

Chiyoung Kim, Min Sung Choi, Jinho Ju, Chanhoe Gu, Donghwan Hwang, Wonseok Choi, Woongsun Jeon, Minhyeok Lee

首次发表
浏览论文内容

中文总结 AI 辅助

ReShoot通过视觉语言模型和视频生成器重新渲染录制的演示,生成视觉多样数据,提升模仿学习策略的鲁棒性,在LIBERO和物理机器人上验证了有效性。

中文摘要 AI 辅助

模仿学习的机器人策略常常过度拟合其训练演示中的视觉条件。因此,物体颜色或背景外观的变化往往导致显著的性能下降。一种常见的缓解策略是在每种新的视觉情境中获取额外的演示;然而,这种方法资源密集,需要对每种要覆盖的外观条件重复访问机器人、受控环境和人工操作。我们引入了ReShoot,一个通过在外观改变的情况下重新渲染先前录制的演示来合成视觉多样性的框架,从而将负担从数据收集转移到生成。一个视觉语言模型为场景生成描述,编辑一个目标属性(例如,背景、物体颜色或材质),一个边缘条件的视频生成器重新渲染两个相机视图以匹配。指令相应更新。动作序列和本体感觉轨迹被原样复制,无需重新标记,因此每个生成的片段保留录制的动作和本体感觉标签。在LIBERO上,使用录制和重新渲染的演示等量混合训练的策略与仅使用录制训练的性能相匹配(96.5%对96.9%)。此外,混合训练集提高了对LIBERO-Plus场景扰动的鲁棒性(85.5%对82.3%)。在两个物理机器人平台上,使用43个和100个预先收集的演示部署ReShoot,将重新着色物体上的成功率分别从0.0%提高到42.9%和47.5%,同时在原始录制外观下保持性能。

英文摘要

Imitation-learned robot policies are frequently overfit to the visual conditions present in their training demonstrations. Consequently, variations in object color or background appearance often induce substantial performance degradation. A common mitigation strategy is to acquire additional demonstrations in each novel visual context; however, this approach is resource-intensive, requiring repeated access to a robot, a controlled environment, and human operation for every appearance condition to be covered. We introduce ReShoot, a framework that synthesizes visual diversity by re-rendering previously recorded demonstrations under altered appearances, thereby shifting the burden from data collection to generation. A vision-language model captions the scene, edits a targeted attribute (e.g., background, object color, or material), and an edge-conditioned video generator re-renders both camera views to match. The instruction is updated accordingly. The action sequence and proprioceptive trajectory are copied verbatim without relabeling, so each generated episode retains the recorded action and proprioceptive labels. On LIBERO, a policy trained on an equal mixture of recorded and re-rendered demonstrations matches the performance of recorded-only training (96.5% vs. 96.9%). Moreover, the mixed training set improves robustness to scene perturbations on LIBERO-Plus (85.5% vs. 82.3%). Across two physical robotic platforms, deploying ReShoot with 43 and 100 pre-collected demonstrations increased the success rate on recolored objects from 0.0% to 42.9% and 47.5%, respectively, while maintaining performance under the original recorded appearance.

发表机构

  • Chung-Ang University(中央大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑