从随机探索中学习规划
Learning to Plan from Random Exploration
- UCLA(加州大学洛杉矶分校)
- University of Minnesota(明尼苏达大学)
- Amazon AGI(亚马逊AGI)
- Salesforce Research(赛富时研究院)
- Lambda(Lambda公司)
- Rice University(莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出利用随机探索数据,通过条件能量模型学习时间关系,实现无需策略改进训练的长距离规划,并在迷宫导航和操作规划中验证了多尺度认知地图特性。
中文摘要 AI 辅助
随机探索在目标指定之前揭示了环境如何被遍历。这种经验能否在不进行策略改进训练的情况下支持长距离规划?我们的随机游走分析解释了时间关系包含的内容:短视界在扩散极限中揭示测地线几何,而较长的视界在混合消除这些区别之前揭示区域之间的连通性。我们通过一个条件能量模型学习这些关系,该模型通过视界条件嵌入估计时间对数密度比。该模型通过噪声对比估计在观测对上训练,无需动作或奖励标签。规划器在向目标移动时在不同视界查询这些学习到的关系。在测试时,一个独立的局部动力学模型预测候选动作结果,时间模型通过选择或聚合跨视界的估计改进来评估它们朝向目标的进展。智能体执行一个动作并使用两个固定模型重新规划。实验证明了从随机探索中使用状态和图像进行长距离迷宫规划。学习到的分数场、嵌入探针和规划路线表现出多尺度认知地图的特性。我们进一步展示了从随机探索中的自我中心导航和从次优数据中的操作规划。
英文摘要
Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the goal. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons. The agent executes one action and replans with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data.