arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29555cs.CVcs.RO

导航世界模型的视觉表示与历史建模

Visual Representation and History Modeling for Navigation World Models

  • Clemson University(克莱姆森大学)
  • University of Ottawa(渥太华大学)

机构由 AI 辅助整理,请以论文原文为准。

Guangfu Guo, Xiaoqian Lu, Rui Liu, Yutong Chen, Kunpeng Liu, Long Cheng

AI总结:

针对导航世界模型在视觉表示选择与历史建模上的两大挑战,提出统一条件流变换器框架,比较五种冻结表示并设计Cached-Linear与平衡门控Delta网络,实现高效长上下文与多查询预测。

AI中文摘要:

导航世界模型(NWMs)预测以行动为条件的视觉未来以用于规划。其设计面临两个实际挑战:选择合适的视觉表示,以及高效建模观测历史以支持重复的候选查询。标准的全局Softmax注意力提供了灵活的交互,但会重复处理相同的历史,导致在长上下文和多查询规划中计算和内存成本不断增加。我们在一个统一的条件流变换器框架内研究这两个问题。我们首先在相同的动力学模型和评估下比较五种冻结的视觉表示。为了减少冗余的历史计算,我们设计了Cached-Linear,一种混合架构,结合局部和移位窗口注意力用于目标混合,以及线性注意力用于可重用的历史访问。我们进一步开发了平衡门控Delta网络(GDN),通过帧级循环记忆增强该设计以进行时间历史建模。在RECON、SACSoN和SCAND上的实验表明,表示选择取决于预测目标:PAE-L在重建上表现最佳,RAE-B在直接预测上最佳,V-JEPA在长时程展开上最佳。在共享历史工作负载下,与全局Softmax相比,Cached-Linear大幅减少了计算和内存,而平衡GDN通过高效的上下文复用改进了选定的直接预测端点。总体而言,我们系统地研究了NWMs的视觉表示和历史建模,并开发了混合可重用历史架构以实现高效的长上下文和多查询预测。

英文摘要:

Navigation World Models (NWMs) predict action-conditioned visual futures for planning. Two practical challenges are central to their design: selecting a suitable visual representation and efficiently modeling observation history for repeated candidate queries. Standard Global-Softmax attention provides flexible interactions but repeatedly processes the same history, leading to increasing computation and memory costs for long contexts and multi-query planning. We study both problems within a unified conditional flow-transformer framework. We first compare five frozen visual representations under the same dynamics model and evaluation. To reduce redundant history computation, we design Cached-Linear, a hybrid architecture that combines local and shifted-window attention for target mixing with linear attention for reusable history access. We further develop Balanced Gated Delta Network (GDN), which augments this design with frame-wise recurrent memory for temporal history modeling. Experiments on RECON, SACSoN, and SCAND show that representation choice depends on the prediction objective: PAE-L performs best for reconstruction, RAE-B for direct prediction, and V-JEPA for long-horizon rollout. Under shared-history workloads, Cached-Linear substantially reduces computation and memory compared with Global-Softmax, while Balanced GDN improves selected direct-prediction endpoints with efficient context reuse. Overall, we systematically study visual representation and history modeling for NWMs and develop hybrid reusable-history architectures for efficient long-context and multi-query prediction.

↑