CosmosAlign:适配世界基础模型用于生成式交通视频预测
CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting
- Simon Fraser University(西蒙菲莎大学)
- Institut Polytechnique de Paris(巴黎综合理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出基于Cosmos3-Nano的CosmosAlign框架,通过两阶段LoRA适配与推理优化,在AI City Challenge 2026 Track 5基准中获76.49分排名第一,实现了高质量交通视频预测。
AI中文摘要:
生成式交通视频预测旨在从短序列观测历史和文本描述中合成长时序、时间上连贯的交通场景未来视频。本文提出了CosmosAlign,一个基于预训练世界基础模型Cosmos3-Nano构建的生成式交通视频预测框架。我们的方法源于一项观察:成功将大型预训练世界模型适配到下游预测任务主要依赖分布对齐,而非提升模型容量。为此,我们提出了两阶段LoRA适配策略:第一阶段将条件模式分布与目标预测任务对齐,第二阶段通过基于大语言模型(LLM)的重描述管道,将训练用描述文本与模型原生结构化提示接口对齐。推理阶段,我们采用完全无需训练的过程进一步提升预测质量,该过程包含基于共识的medoid样本选择,以及静态场景区域的运动自适应融合。CosmosAlign在AI City Challenge 2026 Track 5基准测试中取得了76.49的最终分数,在最终排行榜上排名第一。我们的代码可在this https URL获取。
英文摘要:
Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation history and textual descriptions. In this paper, we present CosmosAlign, a generative traffic video forecasting framework built upon the pretrained Cosmos3-Nano world foundation model. Our approach is motivated by the observation that successfully adapting large pretrained world models to downstream forecasting tasks depends primarily on distribution alignment rather than increased model capacity. To this end, we propose a two-stage LoRA adaptation strategy that first aligns the conditioning-mode distribution with the target forecasting task, and then aligns the training captions with the model's native structured prompting interface through an LLM-based re-captioning pipeline. During inference, we further improve prediction quality using a fully training-free procedure consisting of consensus-based medoid sample selection and motion-adaptive blending of static scene regions. CosmosAlign achieves a final score of 76.49 on the AI City Challenge 2026 Track 5 benchmark, ranking first on the final leaderboard. Our code is publicly available at https://quangminhdinh.github.io/CosmosAlign/.