Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
机构 * School of Computer Science and Engineering, and the MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University(计算机科学与工程学院,以及教育部先进计算机应用技术工程研究中心,北京航空航天大学) ; School of Computer Science and Engineering, the School of Economics and Management, and the MIIT Key Laboratory of Data Intelligence and Management, Beihang University(计算机科学与工程学院,经济管理学院,以及工信部数据智能与管理重点实验室,北京航空航天大学) ; Gaoling School of Artificial Intelligence, Renmin University of China(中关村人工智能学院,中国人民大学) ; DiDi Global Inc.(滴滴出行公司)