机构
*
ByteDance(字节跳动)
;
School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院)
;
National University of Singapore(国立新加坡大学)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
通过融合来自LLM的世界知识与视觉基础模型进行视频事件推理与预测
L'ea Dubois, Klaus Schmidt, Chengyu Wang, Ji-Hoon Park, Lin Wang, Santiago Munoz
机构
*
INRIA(法国国家信息与自动化研究所)
;
Max Planck Institute for Intelligent Systems(人工智能研究所)
;
San Francisco State University(旧金山州立大学)
;
Seoul AI Institute (SAII)(首尔人工智能研究所)
;
Vision & Robotics Center, Tsinghua University(清华大学视觉与机器人中心)
;
Polytechnic University of Madrid(马德里理工大学)
机构
*
School of Software Technology, Zhejiang University, Hangzhou, China(浙江大学软件技术学院)
;
School of Computer Science, Zhejiang University, Hangzhou, China(浙江大学计算机科学学院)
;
School of Software Engineering, Xidian University, Xi’an, China(西安电子科技大学软件工程学院)
CommentsAccepted to AAAI 2026. This arXiv version corresponds to the camera-ready manuscript and includes expanded appendices. Please cite the AAAI 2026 version when available