From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心、自动化研究所、中国科学院) ; Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国) ; School of Artificial Intelligence, University of Chinese Academy of Science, Beijing, China(人工智能学院、中国科学院大学、北京中国) ; Wuhan AI Research, Wuhan, China(武汉人工智能研究、武汉中国) ; MAPLE Lab, Westlake University(MAPLE实验室、西湖大学)
专题命中 视频生成 :video generation(title,abstract);分类 cs.CV