Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
通过融合来自LLM的世界知识与视觉基础模型进行视频事件推理与预测
机构 * INRIA(法国国家信息与自动化研究所) ; Max Planck Institute for Intelligent Systems(人工智能研究所) ; San Francisco State University(旧金山州立大学) ; Seoul AI Institute (SAII)(首尔人工智能研究所) ; Vision & Robotics Center, Tsinghua University(清华大学视觉与机器人中心) ; Polytechnic University of Madrid(马德里理工大学)
专题命中 其他安全 :alignment(abstract)
AI总结 本文提出融合视觉基础模型与LLM的世界知识,以提升视频事件推理与预测能力,实现从简单识别到高级认知理解的突破。
Comments 22 pages, 4 figures