发表机构
Pantheon Industries(万神殿工业公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出RP1方法,将良好搜索规则强化到神经规划器中,可改进多步计划,在多个机器人任务中性能优于手工搜索算法,且效率更高。
AI 中文摘要
人类通过构建计划并利用内部世界模型在脑中模拟其结果来解决复杂问题。机器学习已生成能类似预测动作序列结果的世界模型,但候选计划的改进尚未被完全学习。当前规划器要么是手工设计的,要么是从手工设计的优化器中蒸馏而来,要么仅被学习用于告知摊销策略而非修改计划本身。我们提出了强化规划(Reinforced Planning)方法,其核心思想是通过将良好的搜索规则强化到神经规划器中来学习搜索。我们的实现RP1既学习如何通过评判器(critic)评估想象的结果,又学习如何通过完全从想象的世界模型回推(roll-out)中离线训练的优化器来改进多步计划。据我们所知,RP1是首个完全学习如何改进多步计划的方法,此外,它可独立于任何预训练的潜世界模型进行训练并附加到其上。在视觉导航、机械臂抓取和机器人操作任务中,使用两个世界模型主干时,RP1的性能显著优于手工设计的搜索算法,在多个设置中达到近乎完美的成功率,同时使用的世界模型回推量减少了1000倍,且在并行规划器推理下比最强替代方案快达67倍。
英文摘要
Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has made substantial progress in learning world models that predict the consequences of action sequences, yet the procedures used to plan with these models remain largely hand-designed. Most planners rely on fixed search or optimization rules; approaches that learn aspects of search typically imitate a predefined optimizer or use planning to inform an amortized policy, rather improving multi-step plans. We introduce \textbf{Reinforced Planning}, a method that learns the plan-update itself by reinforcing update rules that produce better plans, using gradients propagated through a differentiable world model. We instantiate Reinforced Planning in RP1, which learns a critic over imagined outcomes via temporal-difference learning and a neural plan-improvement operator trained via imagined rollouts with a pretrained world model. RP1 can be trained fully offline without environment interaction; environment episodes are used only for checkpoint selection. Across visual navigation, arm reaching, and robotic manipulation on two world-model backbones, RP1 matches or exceeds existing planners, achieving near-perfect success in several settings while using $1{,}000\times$ fewer world-model rollouts than the strongest alternative (CEM) and planning up to $67\times$ faster under concurrent planners inference.
CommentsPreprint