发表机构
University of British Columbia; Inverted AI; Amii(不列颠哥伦比亚大学; Inverted AI; Amii)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出镜像学习框架,通过视频扩散模型的视角变换与逆动力学模型合成镜像数据,用其增强行为克隆训练可提升策略性能,为数据收集提供替代方案。
AI 中文摘要
我们从第三人称观察的视角研究模仿学习,提出镜像学习框架:从被动观察中获取可执行策略。行为克隆(BC)在密集、对齐良好的第一人称数据下表现出色,但无法利用人类和动物常利用的第三人称演示带来的丰富观察信号。我们引入一种方法,由两部分组成:(i)利用微调后的视频扩散模型学习视角变换,使学习者处于演示者的视角;(ii)逆动力学模型,在学习者的控制空间中推断动作轨迹。这能合成镜像数据,即从演示者行为的第三人称观察生成的伪第一人称专家数据。实证表明,仅镜像数据就能训练出有效策略,用镜像数据增强第一人称BC训练可进一步提升下游策略性能。我们的结果表明,现代生成式世界模型隐含编码了足够的结构,可作为依赖远程操作的数据收集的可扩展且安全的替代方案。
英文摘要
We investigate imitation learning through the lens of third-person observation and propose a framework for mirror learning: acquiring actionable policies from passive observation. While behavior cloning (BC) excels under dense, well-aligned first-person data, it fundamentally fails to leverage the rich observational signals arising from third-person demonstrations that humans and animals routinely exploit. We introduce a method that composes (i) a learned perspective transformation that places learners in demonstrators' shoes using a fine-tuned video diffusion model and (ii) an inverse dynamics model that infers action trajectories in the learners' control space. This enables the synthesis of mirror data, pseudo first-person expert data generated from third-person observations of demonstrator behavior. Empirically, we show that mirror data alone can train effective policies, and that augmenting first-person BC training with mirror data further improves downstream policy performance. Our results suggest that modern generative world models implicitly encode sufficient structure to enable a scalable and safe alternative to teleoperation-heavy data collection.