机构
*
Nankai University(南开大学)
;
Northwestern Polytechnical University(西北工业大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
NKIARI
机构
*
School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Cloud and AI BU, Huawei(华为云与AI业务部)
;
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
Hi-GaTA: Hierarchical Gated Temporal Aggregation Adapter for Surgical Video Report Generation
Hi-GaTA:用于外科视频报告生成的分层门控时间聚合适配器
Kedi Sun, Chaohui Dang, Yue Feng, James Glasbey, Theodoros N. Arvanitis, Le Zhang
机构
*
School of Engineering, College of Engineering and Physical Sciences, University of Birmingham, Birmingham, UK(英国伯明翰大学工程学院)
;
School of Computer Science, University of Birmingham, Birmingham, UK(英国伯明翰大学计算机科学学院)
;
Department of Applied Health Sciences, University of Birmingham, Birmingham, UK(英国伯明翰大学应用健康科学系)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
GraphThinker: 通过事件图思维强化时间感知的视频推理
Zixu Cheng, Da Li, Jian Hu, Yuhang Zang, Ziquan Liu, Shaogang Gong, Wei Li
机构
*
Queen Mary University of London(伦敦玛丽女王大学)
;
Samsung AI Centre Cambridge(剑桥三星人工智能中心)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Nanyang Technological University(南洋理工大学)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Kling Team, Kuaishou Technology(快手科技 Kling 团队)
;
Institute of Software Chinese Academy of Sciences(中国科学院软件研究所)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
基于检测器的视频大语言模型用于高效的时空定位
Shida Gao, Feng Xue, Xiangfeng Wang, Anlong Ming, Zhaowen Lin, Haiyang Zhang, Teng Long, Nicu Sebe, Yihua Shao, Haozhe Wang, Wei Wang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Trento(特伦特大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Hong Kong University of Science and Technology(香港科技大学)
;
ZTE Corporation(中兴通讯)
All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction
所有变化可能都有不变原则:通过设计概念再现改进永动有害迷因检测
Ziyou Jiang, Mingyang Li, Junjie Wang, Yuekai Huang, Jie Huang, Zhiyuan Chang, Zhaoyang Li, Qing Wang
机构
*
State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室)
;
Science and Technology on Integrated Information System Laboratory Institute of Software Chinese Academy of Sciences(软件研究所信息集成系统技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
机构
*
Institute for AI, Peking University(人工智能研究院,北京大学)
;
Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院)
;
School of Psychological and Cognitive Sciences, Peking University(心理学与认知科学学院,北京大学)
;
School of Computer Science, Peking University(计算机科学学院,北京大学)
;
Yuanpei College, Peking University(元培学院,北京大学)
;
School of Foreign Languages, Peking University(外语学院,北京大学)
;
School of EECS, Peking University(电子工程与科学学院,北京大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
State Key Lab of General AI(通用人工智能国家重点实验室)
;
Nat’l Eng. Research Center of Visual Technology(视觉技术国家工程研究中心)
;
Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京行为与心理健康重点实验室,北京大学)
;
Embodied Intelligence Lab, PKU-Wuhan Institute for Artificial Intelligence(具身智能实验室,北京大学-武汉人工智能研究院)
Bisecle: Binding and Separation in Continual Learning for Video Language Understanding
Yue Tan, Xiaoqian Hu, Hao Xue, Celso De Melo, Flora D. Salim
机构
*
School of Computer Science University of New South Wales(计算机科学学院新南威尔士大学)
;
School of Computer Science and Engineering University of New South Wales(计算机科学与工程学院新南威尔士大学)
;
DEVCOM Army Research Laboratory(陆军研究实验室)
专题命中
视频多模态
:multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV
机构
*
Central South University(中南大学)
;
Tsinghua University(清华大学)
;
South China Normal University(华南师范大学)
;
ByteDance Inc(字节跳动公司)
;
University of Zaragoza(阿拉维达大学)
;
CosmosMind
;
Wuhan University(武汉大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Southeast University(东南大学)
;
Tencent(腾讯公司)
;
Nankai University(南开大学)
;
Supermicro Computer Inc(Supermicro计算机公司)
;
Huazhong University of Science and Technology(华中科技大学)
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering
MemoryCard: 面向长视频问答的主题感知多模态线索压缩
Qing Yang, Pengcheng Huang, Xinze Li, Zhenghao Liu, Yukun Yan, Yu Gu, Ge Yu, Gang Li, Maosong Sun
机构
*
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Digital China Group(数字中国集团)