JEPA引导扩散:用于生成式交通预测的预测性视觉-语言条件化
JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecasting
浏览论文内容
中文总结 AI 辅助
提出一种解耦的交通预测框架,利用冻结的V-JEPA编码器提取预测性潜在表示,通过轻量级对齐模块引导冻结的Cosmos扩散模块生成未来视频,在AI City Challenge 2026 Track 5上以75.1297分排名第三,显著降低训练成本。
中文摘要 AI 辅助
准确的交通预测既需要理解场景动态,又需要合成逼真的未来观测。最近的基于扩散的视频生成模型能产生视觉上合理的预测,但需要昂贵的端到端训练,且常常将场景理解与图像合成纠缠在一起。在这项工作中,我们提出了一种解耦的预测框架,将未来表示学习与视频生成分离。一个冻结的V-JEPA编码器首先从观测到的交通视频中提取预测性潜在表示,在语义潜在空间中捕获底层场景动态。然后,一个轻量级的潜在对齐模块将这些表示投影到冻结的Cosmos扩散模块的条件空间中,从而无需重新训练大型生成模型即可实现未来视频合成。通过冻结所有基础模型并仅训练轻量级的对齐模块,所提出的框架大幅降低了优化复杂度,同时保持了预测能力。在AI City Challenge 2026 Track 5基准上的实验结果表明,所提出的方法取得了75.1297的分数,在竞赛中排名第三。这些结果表明,V-JEPA学习到的预测性世界表示能有效引导下游视频生成,为端到端扩散式预测提供了一种实用且高效的替代方案。
英文摘要
Accurate traffic forecasting requires both understanding scene dynamics and synthesizing realistic future observations. Recent diffusion-based video generation models produce visually plausible predictions but require expensive end-to-end training and often entangle scene understanding with image synthesis. In this work, we propose a decoupled forecasting framework that separates future representation learning from video generation. A frozen V-JEPA encoder first extracts predictive latent representations from the observed traffic videos, capturing the underlying scene dynamics in a semantic latent space. A lightweight latent alignment module then projects these representations into the conditioning space of a frozen Cosmos diffusion module, enabling future video synthesis without retraining the large generative model. By freezing all foundation models and training only the lightweight alignment module, the proposed framework substantially reduces optimization complexity while preserving forecasting capability. Experimental results on the AI City Challenge 2026 Track 5 benchmark demonstrate that the proposed method achieved a score of 75.1297, ranking third in the competition. These results suggest that predictive world representations learned by V-JEPA can effectively guide downstream video generation, providing a practical and efficient alternative to end-to-end diffusion-based forecasting.
发表机构
- Ho Chi Minh City University of Technology and Engineering(胡志明市技术教育大学)
机构由 AI 辅助整理,请以论文原文为准。