发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RECON方法,通过分离主策略与探索器,基于可恢复性条件的不确定性引导探索,提升基于模型模仿学习在运动、导航和操作任务中的交互效率与性能。
AI 中文摘要
基于模型的模仿学习(MBIL)通过在从学习到的世界模型生成的想象轨迹上优化策略,提高了真实环境交互的效率。然而,模型诱导的占用与真实环境占用之间的差距使得策略学习对模型误差敏感。保守的MBIL在策略优化过程中减轻了模型利用,但当真实环境交互由相同的保守策略收集时,专家分布周围的不确定区域仍然采样不足。另一方面,通用的不确定性驱动的探索可能将交互分配给新颖但与任务无关的动态。我们提出了用于基于模型的模仿学习的可恢复性条件探索(RECON)。RECON通过维护一个用于任务执行的主策略和一个用于真实环境交互的探索器,将保守策略学习与主动数据收集分离。探索器基于以从多步主策略想象中估计的可恢复性为条件的认知不确定性进行优化,将数据收集集中在主策略仍能返回专家行为附近的未知状态上。在运动、导航和操作任务上的实验显示,在交互效率、模仿性能和鲁棒性方面的一致提升,表明RECON将真实环境交互引导到专家分布周围先前方法未充分探索的恢复区域,从而学习到更适合模仿的世界模型。
英文摘要
Model-based imitation learning (MBIL) improves real-environment interaction efficiency by optimizing policies on imagined rollouts from a learned world model. However, the gap between model-induced and real-environment occupancies makes policy learning sensitive to model error. Conservative MBIL mitigates model exploitation during policy optimization, but when real-environment interactions are collected by the same conservative policy, uncertain regions around the expert distribution remain insufficiently sampled. Generic uncertainty-driven exploration, on the other hand, may allocate interaction to novel but task-irrelevant dynamics. We propose REcoverability-CONditioned Exploration for Model-Based Imitation Learning (RECON). RECON separates conservative policy learning from active data collection by maintaining a main policy for task execution and an explorer for real-environment interaction. The explorer is optimized based on epistemic uncertainty conditioned on recoverability estimated from multi-step main-policy imagination, focusing data collection on unknown states from which the main policy can still return toward expert behavior. Experiments on locomotion, navigation and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness, indicating that RECON directs real-environment interaction toward recovery regions around the expert distribution that are underexplored by prior methods, and thereby learns a world model better suited for imitation.