DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Alibaba Group(阿里巴巴集团) ; Peking University(北京大学) ; Luoyang Institute for Robot and Intelligent Equipment(洛阳机器人与智能装备研究所)
Comments Accepted by ICCV 2025