发表机构
Technical University of Berlin; Heinrich Heine University Düsseldorf(柏林工业大学; 杜塞尔多夫海因里希·海涅大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究扩展了VideoGAN框架,改进语义表示与轨迹提取方式,系统研究大视场,引入评估框架,实现高效逼真轨迹生成,适配自动驾驶下游任务。
AI 中文摘要
逼真且多样的轨迹生成是实现更高水平车辆自动化的核心。基于规则和经典学习的方法可能难以捕捉交通行为的复杂性,而生成模型已在其他领域证明其可处理同等程度的复杂性。本文基于先前基于生成对抗网络(GAN)的语义鸟瞰图交通生成的研究,在几个关键方面扩展了所提出的框架:改进语义表示,将轨迹提取过程替换为基于图的关联方法,并系统研究越来越大的视场;此外,引入定量评估框架以评估生成视频中的幻觉和物体恒存性。实验表明,该框架可泛化到更大、更复杂的交通场景,同时保持统计上逼真的轨迹以及交通参与者之间连贯的空间关系。在150 GPU小时的训练且对长达20秒的场景推理时间低于20毫秒的情况下,结果表明基于视频的GAN仍是逼真轨迹生成的高效且可扩展方法,即使在大得多的交通场景中也适用,非常适合自动驾驶中的预测、规划和模拟等下游任务。
英文摘要
Realistic and diverse trajectory generation is central to enabling higher levels of vehicle automation. While rule-based and classical learning-based methods may struggle to capture the complexity of traffic behavior, generative models have already demonstrated in other fields that they can handle a comparable level of complexity. In this paper, we build upon previous work on generative adversarial network (GAN)-based semantic bird's-eye-view traffic generation and extend the proposed framework in several key aspects. We improve the semantic representation, replace the trajectory extraction procedure with a graph-based association method, and systematically investigate increasingly larger fields of view. In addition, we introduce a quantitative evaluation framework to assess hallucinations and object permanence in generated videos. Our experiments demonstrate that the framework generalizes to larger and more complex traffic scenes while maintaining statistically realistic trajectories and coherent spatial relationships between traffic participants. Within 150GPU hours of training and with inference times below 20ms for scenes of up to 20s, our results demonstrate that video-based GANs remain an efficient and scalable approach for realistic trajectory generation, even in substantially larger traffic scenes, making them well suited for downstream tasks such as prediction, planning, and simulation in automated driving.