arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Latent-Foresight:端到端学习潜在世界模型的可预测表示

Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis

arXiv 2610.01942首次发表:更新:

发表机构

Archimedes, Athena Research Center; National Technical University of Athens; University of Crete; IACM-Forth(阿基米德,雅典娜研究中心; 雅典国立技术大学; 克里特大学; 希腊研究与技术基金会计算与应用数学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Latent-Foresight端到端框架,联合学习潜在分词器和流式生成动力学模型,塑造可预测表示,在多个未来场景理解任务上优于两阶段基线,消除单独训练阶段。

AI 中文摘要

预测场景的未来演化是世界建模的基本能力。近期工作表明,在视觉基础模型(VFM)的特征空间中操作能够产生语义丰富的表示,支持多样的未来场景理解任务。然而,现有方法依赖于两阶段流程,即首先使用固定降维(如PCA)或独立训练的自编码器压缩VFM特征,然后在得到的冻结潜在空间之上训练一个单独的预测器。这种表示学习与时间预测之间的解耦,以及直接在原始VFM特征上应用预测器的方法,无法保证潜在空间被结构化以支持可预测的动态。在这项工作中,我们提出了Latent-Foresight,一个端到端框架,它联合学习一个潜在分词器和一个基于流的生成动力学模型,明确塑造表示以支持时间可预测性。为了实现稳定的联合优化,我们引入了几个关键设计选择,以防止潜在坍缩并使重建与生成目标对齐。大量实验表明,我们的方法学习了更时间连贯的潜在表示,并在多个未来场景理解任务和预测范围上持续优于两阶段基线,同时消除了单独的训练阶段,包括在高分辨率适应期间。我们在https URL提供实现代码和模型权重。

英文摘要

Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of Vision Foundation Models (VFMs) yields semantically rich representations that support diverse future scene understanding tasks. However, existing approaches rely on two-stage pipelines, where VFM features are first compressed using fixed dimensionality reduction (e.g., PCA) or independently trained autoencoders, and a separate predictor is trained on top of the resulting frozen latent space. This decoupling between representation learning and temporal prediction, as well as approaches that apply predictors directly on raw VFM features, provides no guarantee that the latent space is structured for predictable dynamics. In this work, we propose Latent-Foresight, an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model, explicitly shaping the representation to support temporal predictability. To enable stable joint optimization, we introduce several key design choices that prevent latent collapse and align reconstruction with generative objectives. Extensive experiments show that our approach learns more temporally coherent latent representations and consistently outperforms two-stage baselines across multiple future scene understanding tasks and prediction horizons, while eliminating separate training stages, including during high-resolution adaptation. We provide the implementation code and model weights at https://github.com/Sta8is/Latent-Foresight

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑