发表机构
The University of Manchester; University of Technology Sydney; Shandong University; Shanghai Jiao Tong University(曼彻斯特大学; 悉尼科技大学; 山东大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出DF³框架,通过冻结视觉基础模型、MACF机制及专用任务查询,在潜在空间直接预测特征并生成任务输出,在性能与效率上实现平衡,适用于自主导航的世界建模。
AI 中文摘要
从视频序列预测未来状态是自主机器人系统的关键挑战,也是世界建模的核心目标。现有像素级生成方法不可避免地过度强调与任务无关的细节,导致计算开销过高;而基于潜在空间的方法虽试图通过直接预测特征缓解该问题,但仍依赖重型解码器进行状态到任务的映射,成为计算瓶颈。本研究提出无解码器特征预测(DF³)框架,该框架完全在潜在空间中建模世界演化并直接生成任务输出,彻底消除对解码器的需求。具体而言,DF³将可学习空间查询注入冻结视觉基础模型的终端块,以直接提取未来状态表征;通过采用轻量统一的运动感知上下文融合(MACF)机制,该机制将粗光流扭曲与细粒度潜在互相关无缝结合,使这些查询与历史令牌表征交互,明确对齐并预测下一帧的特征;随后,一组专用任务查询探测这些预测特征以完成下游任务。在公共基准上的大量实验及在机器人模拟器中的零样本部署表明,DF³的性能与现有最先进方法相当,同时在集成感知与控制方面提供更优的效率和灵活性。
英文摘要
Forecasting future states from video sequences is a critical challenge for autonomous robotic systems and a fundamental objective of world modeling. Prior generative methods operating at the pixel level inevitably overemphasize task-irrelevant details, leading to prohibitive computational overhead. While latent-based approaches attempt to mitigate this by predicting features directly, the persistent reliance on heavy decoders for state-to-task mapping remains a computational bottleneck. In this work, we propose Decoder-Free Feature Forecasting (DF$^3$), a novel framework that models world evolution entirely within the latent space and directly derives task outputs, completely eliminating the need for a decoder. Specifically, DF$^3$ injects learnable spatial queries into the terminal blocks of a frozen vision foundation model to extract future state representations directly. By employing a lightweight, unified Motion-Aware Context Fusion (MACF) mechanism that seamlessly integrates coarse flow warping with fine-grained latent cross-correlation, these queries interact with historical token representations to explicitly align and forecast the feature of the next frame. Subsequently, a specialized set of task queries probes these forecasted features for the downstream task. Extensive experiments on public benchmarks and zero-shot deployment in a robotic simulator demonstrate that DF$^3$ achieves performance comparable to state-of-the-art methods while offering superior efficiency and flexibility for integrated perception and control.