Video Understanding: From Geometry and Semantics to Unified Models
视频理解:从几何与语义到统一模型
机构 * Department of Computer Science(计算机科学系) ; University of Copenhagen(哥本哈根大学) ; College of Computer Science(计算机科学学院) ; Nankai University(南开大学) ; School of Computer and Communication Sciences(计算机与通信科学学校) ; EPFL(苏黎世联邦理工学院) ; Department of Computer Science & Engineering(计算机科学与工程系) ; Washington University in St. Louis(圣路易斯华盛顿大学) ; Computer Vision Lab(计算机视觉实验室) ; University of Würzburg(乌尔姆大学) ; Computer Science Research Centre(计算机科学研究中心) ; University of Surrey(萨里大学) ; School of Electrical, Computer and Telecommunications Engineering(电气、计算机和电信工程学院) ; University of Wollongong(沃林根大学) ; Language Technology Lab(语言技术实验室) ; University of Cambridge(剑桥大学) ; School of Artificial Intelligence(人工智能学院) ; Beijing Institute of Technology(北京理工大学) ; Institute for Computer Science(计算机科学研究所) ; INSAIT
AI总结 本文综述了视频理解的发展,从低层几何理解到高层语义理解和统一模型,探讨了时间动态和视觉上下文建模的重要性,并总结了当前研究趋势和挑战。
Comments A comprehensive survey of video understanding, spanning low-level geometry, high-level semantics, and unified understanding models
Journal ref Machine Intelligence Research 2026