arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多远而非多近:用于潜在世界模型规划的学习时间度量

How Long, Not How Close: A Learned Temporal Metric for Planning in Latent World Models

Lama Moukheiber, Haotian Xue, Yongxin Chen

arXiv 2610.04988首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对潜在世界模型规划中目标距离较远时排序失效的问题,提出时间距离目标TEMPO,利用现有演示学习状态间步数映射并融入规划代价,在多种环境中显著提升规划性能。

AI 中文摘要

潜在世界模型通过将冻结的预测器在候选动作序列下向前滚动,并根据想象终点状态与目标之间的潜在距离对候选进行排序来进行规划。然而,当目标位于多个计划之外时,这种排序会失效,因为潜在距离衡量的是终点状态与目标的相似程度,而非距离达成目标还有多远。为解决这一问题,我们提出TEMPO,一种时间距离规划目标,它不改变世界模型,仅利用已用于训练模型的演示数据进行学习,并且为规划器的搜索增加可忽略不计的代价。TEMPO学习一个冻结潜在空间的小型映射,其中一段情节中两个状态之间的距离反映它们之间的环境步数,并将该距离融入规划器的代价中。它不需要奖励、策略或成功标签,并且作为一种代价而非模型,适用于每个状态具有一个潜在向量且通过潜在距离进行规划的冻结世界模型。我们在十一个模拟环境(例如迷宫导航、桌面推挤、机械臂控制和三维操作)中使用LeWM和PLDM规划器评估TEMPO。借助一个最多为一次计划的算术增加0.3%的小型MLP,TEMPO在每种目标距离下都改进了两个规划器,包括其评估中的单计划设置,将LeWM在TwoRoom中距目标三个计划时的成功率从36%提升至99%,并在广泛的2D和3D导航、到达和操作任务中保持竞争力。

英文摘要

Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans away, because the latent distance measures how closely an end state resembles the goal rather than how far it remains from reaching it. To address this, we propose TEMPO, a temporal-distance planning objective that leaves the world model untouched, learns only from the recorded trajectories already used to train it, and adds negligible cost to the planner's search. TEMPO learns a small map of the frozen latent in which the distance between two states of an episode reflects the number of environment steps between them, and blends this distance into the planner's cost. It requires no rewards, policies or success labels and, being a cost rather than a model, applies to frozen world models with one latent vector per state that plan by a latent distance. We evaluate TEMPO on eleven simulated environments (e.g., maze navigation, tabletop pushing, robotic arm control and three-dimensional manipulation) with the LeWM and PLDM planners. With a small MLP that adds at most 0.3% to a plan's arithmetic, TEMPO improves both planners at every goal distance, including the one-plan setting of their evaluations, raises LeWM from 36% to 99% on TwoRoom three plans from the goal, and remains competitive on a broad range of 2D and 3D navigation, reaching and manipulation tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑