IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
机构 * University of Science and Technology of China(中国科学技术大学) ; Hefei Institutes of Physical Science,Chinese Academy of Sciences(合肥物理研究所,中国科学院) ; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究所,合肥综合性国家科学中心) ; State Key Lab. for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) ; University of Chicago(芝加哥大学) ; National University of Defense Technology(国防科技大学) ; Harbin Institute of Technology(哈尔滨工业大学)
专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV