发表机构
Tsinghua University; Tencent Hunyuan; Zhongguancun Academy(清华大学; 腾讯混元; 中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CoDrive提出跨车辆多视角视频生成框架,通过共享世界坐标与混合训练实现精确轨迹控制下的多车观测一致性,提升轨迹可控性和几何、实例一致性。
AI 中文摘要
现实世界的驾驶本质上是一个多智能体过程,然而现有的大多数驾驶世界模型仅从单一自车视角生成观测。将这些模型独立扩展到多车辆场景,并不能确保不同智能体观察到一致的共享世界。我们提出了CoDrive,一个跨车辆、多视角的驾驶视频生成框架,能够以精确的相机轨迹控制,联合生成共享同一动态场景的车辆的观测。CoDrive将局部自注意力(用于建模每辆车各视角间的时空依赖)与全局自注意力(用于实现车辆间的信息交换和一致性建模)交错结合。为了显式编码车辆间的空间关系,所有相机轨迹均表示在共享的世界坐标系中,并通过投影相对位置编码注入注意力层。我们进一步采用了渐进式混合任务训练策略,将大规模真实世界单智能体数据与合成跨智能体交互数据相结合,使模型既能受益于真实世界的外观分布,又能从仿真中学习跨智能体一致性。为了进行系统性评估,我们引入了CoDrive-Bench基准,涵盖真实和合成的多车辆场景,并评估轨迹可控性、场景几何一致性和实例级一致性。实验表明,CoDrive在保持具有竞争力的视觉质量的同时,提升了轨迹可控性以及跨智能体的几何和实例一致性。
英文摘要
Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video generation framework that jointly generates observations of vehicles sharing the same dynamic scene with precise camera-trajectory control. CoDrive interleaves local self-attention, which models spatiotemporal dependencies among the views of each vehicle, with global self-attention, which enables information exchange and consistency modeling across vehicles. To explicitly encode their spatial relationships, all camera trajectories are represented in a shared world coordinate system and injected into the attention layers through projective relative positional encoding. We further adopt a progressive mixed-task training strategy that combines large-scale real-world single-agent data with synthetic cross-agent interaction data, allowing the model to benefit from real-world appearance distributions while learning cross-agent consistency from simulation. For systematic evaluation, we introduce CoDrive-Bench, a benchmark covering real and synthetic multi-vehicle scenarios and evaluating trajectory controllability, scene geometry consistency, and instance-level consistency. Experiments show that CoDrive improves trajectory controllability and cross-agent geometric and instance consistency while maintaining competitive visual quality.
Comments28 pages, 6 figures