视觉模仿学习中视点泛化策略关键因素的实证研究
An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning
浏览论文内容
中文总结 AI 辅助
该研究通过受控实验发现,保留密集视觉标记并让动作头参与几何推理能显著提升视觉模仿学习策略的跨视点泛化能力,并在模拟和真实世界随机相机配置下验证了有效性。
中文摘要 AI 辅助
视觉模仿学习是训练机器人操作策略的一种有前景的方法,能够完成多种多样的任务。然而,目前的策略对视点扰动仍然脆弱,这使得在多样化环境中的部署面临挑战。我们进行了一项受控的实证研究,探讨哪些设计选择能使视觉运动策略跨视点泛化。我们发现,当保留密集的视觉标记并且动作头参与几何推理时,视点泛化能力会得到提升。在一套涵盖广泛相机位姿的模拟任务套件上,我们表明这些设计选择所产生的策略在跨视点情况下仍保持性能。作为实际结果,采用这些设计选择训练的策略还能在随机相机配置下从模拟零样本迁移到现实世界。
英文摘要
Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choices allow visuomotor policies to generalize across viewpoints. We find that viewpoint generalization improves when dense visual tokens are retained and the action head participates in geometric reasoning. On a suite of simulated tasks that span a wide range of camera poses, we show that these design choices yield a policy that remains performant across viewpoints. As a practical consequence, a policy trained with these design choices also transfers zero-shot from simulation to the real world under random camera configurations.