arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RIWANav:用于城市导航的具有自我改进能力的递归世界-动作模型

RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation

Jing Xie, Shouwei Ruan, Yubin Wang, Yuxiang Zhang, Haitao Yang, Songchang Jin, Dianxi Shi

arXiv 2610.08640首次发表:更新:

发表机构

Shanghai Jiao Tong University; Beihang University; The Hong Kong University of Science and Technology; Tsinghua University; The University of Texas at Austin; Intelligent Game and Decision Lab (IGDL)(上海交通大学; 北京航空航天大学; 香港科技大学; 清华大学; 德克萨斯大学奥斯汀分校; 智能博弈与决策实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RIWANAV提出递归自我改进框架,通过交替更新世界模型与策略,利用想象反馈和自课程学习提升城市导航性能,优于现有方法。

AI 中文摘要

长时程城市导航需要一系列连续的局部决策,而这些决策的误差会随时间累积。模仿学习(IL)很少能从失败中学习,而物理试错强化学习(RL)成本高昂。动作条件世界模型可以通过预测候选动作的视觉后果来提供想象反馈。然而,随着策略的演进,冻结的世界模型可能变得不太可靠。在本文中,我们介绍了RIWANAV,一个后训练框架,它将世界模型和动作模型(策略)的耦合适应视为任务特定的递归自我改进(RSI)。每个循环交替进行两次更新。世界模型通过想象的结果评估策略动作,为群体相对策略优化(GRPO)提供比较反馈。改进后的策略随后构建一个基于地面的自我课程,通过行为新颖性和预测误差选择与专家一致的动作-视频对。精炼后的世界模型为下一次策略更新提供反馈,从而闭合递归自我改进循环。实验表明,RIWANAV优于训练基线和先前方法,验证了所提出的策略与世界模型之间的递归自我改进循环。真实世界试验进一步证明了其实用适用性。

英文摘要

Long-horizon urban navigation requires sequential local decisions whose errors can compound over time. Imitation learning (IL) rarely learns from failures, while physical trial-and-error reinforcement learning (RL) is costly. Action-conditioned world models can provide imagined feedback by predicting visual consequences for candidate actions. However, a frozen world model may become less reliable as the policy evolves. In this paper, we introduce RIWANAV, a post-training framework that casts the coupled adaptation of a world model and an action model (policy) as task-specific recursive self-improvement (RSI). Each cycle alternates two updates. The world model evaluates policy actions through imagined outcomes, providing comparative feedback for group-relative policy optimization (GRPO). The improved policy then constructs a grounded self-curriculum, selecting expert-consistent action-video pairs by behavioral novelty and prediction error. The refined world model supplies feedback for the next policy update, closing the recursive self-improvement loop. Experiments show that RIWANAV outperforms training baselines and prior methods, validating the proposed recursive self-improvement loop between the policy and world model. Real-world trials further demonstrate its practical applicability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑