Enhancing Large Vision Model in Street Scene Semantic Understanding through Leveraging Posterior Optimization Trajectory
Wei-Bin Kou, Qingfeng Lin, Ming Tang, Jingreng Lei, Shuai Wang, Rongguang Ye, Guangxu Zhu, Yik-Chung Wu
机构
*
Department of Electrical and Electronic Engineering, The University of Hong Kong(香港大学电子与电气工程系)
;
Shenzhen Research Institute of Big Data(大数据研究深圳研究所)
;
Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
Anthony Chen, Wenzhao Zheng, Yida Wang, Xueyang Zhang, Kun Zhan, Peng Jia, Kurt Keutzer, Shanghang Zhang
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
Li Auto Inc.(利亚德公司)
;
UC Berkeley(加州大学伯克利分校)
Knowledge-Informed Multi-Agent Trajectory Prediction at Signalized Intersections for Infrastructure-to-Everything
Huilin Yin, Yangwenhui Xu, Jiaxiang Li, Hao Zhang, Gerhard Rigoll
机构
*
College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院)
;
Shanghai Institute of Intelligent Science and Technology, Tongji University(同济大学上海智能科学与技术研究院)
;
Department of Electrical and Computer Engineering, Technical University of Munich(慕尼黑技术大学电子与计算机工程系)
CommentsWe are excited to announce that this paper has been accepted for oral presentation at the AAAI 2025 Main Conference. We are grateful for the insightful feedback from the reviewers and look forward to contributing to the discussions at AAAI
Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles