机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Nanyang Technological University(南洋理工大学)
;
Qwen Team, Alibaba Group(阿里巴巴集团Qwen团队)
机构
*
institutetext: MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models Yue Wu Changyuan Wang Zixuan Wang Shilin Ma Yansong Tang(机构文本:MorphoQuant:多模态大语言模型的模态感知量化 Yue Wu 王昌元 王梓轩 马世林 唐彦松)
DeepIPCv3: Event-Aware Multi-Modal Sensor Fusion for Sudden Pedestrian Crossing Avoidance
DeepIPCv3: 面向突发行人穿越避让的事件感知多模态传感器融合
Oskar Natan, Andi Dharmawan, Aufaclav Zatu Kusuma Frisky, Jazi Eko Istiyanto, Jun Miura
机构
*
Department of Computer Science and Electronics, Universitas Gadjah Mada(计算机科学与电子系,加雅马达大学)
;
Department of Computer Science and Engineering, Toyohashi University of Technology(计算机科学与工程系,东福士大学)
机构
*
Multimodal Intelligence Lab(多模态智能实验室)
;
Department of Computer Science(计算机科学系)
;
University of Exeter(埃克塞特大学)
;
School of Computer Science(计算机科学学院)
;
University of Leeds(利兹大学)
;
School of Computer Science and Informatics(计算机科学与信息学学院)
;
University of Liverpool(利物浦大学)
;
University of Birmingham(伯明翰大学)
;
Machine Intelligence + x Group(机器智能+X小组)
机构
*
Department of Logistics and Maritime Studies, the Hong Kong Polytechnic University(物流及海运研究系,香港理工大学)
;
Research Centre for ESG Advancement (RCESGA), the Hong Kong Polytechnic University(ESG进步研究中心(RCESGA),香港理工大学)
;
School of Navigation, Wuhan University of Technology(航海学院,武汉理工大学)
CAM-VFD: Cross-Attention Multimodal Video Forgery Detection
CAM-VFD: 跨注意力多模态视频伪造检测
Hoda Osama Elkhodary, Sherin Mostafa Youssef, Marwa Elshenawy, Dalia Sobhy
机构
*
Computer Engineering Department, College of Engineering and Technology, Arab Academy for Science, Technology and Maritime Transport(计算机工程系,工程与技术学院,阿拉伯科学、技术与海运交通学院)
Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model
基于多模态大语言模型的遥感活动检测的时空感知
David F. Ramirez, Tim Overman, Kristen Jaskie, Andreas Spanias
机构
*
SenSIP Center, School of ECEE, Arizona State University(SenSIP中心,电子与计算机工程学院,亚利桑那州立大学)
;
Prime Solutions Group Inc(Prime Solutions Group公司)
;
Intelligence Advanced Research Projects Activity(智能高级研究计划局)
DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
DeepSport: 一种基于代理强化学习的多模态大语言模型,用于通过主动推理实现综合体育视频理解
Junbo Zou, Haotian Xia, Zhen Ye, Shengjie Zhang, Christopher Lai, Vicente Ordonez, Weining Shen, Hanjie Chen
机构
*
Georgia Institute of Technology(佐治亚理工学院)
;
Rice University(Rice大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
University of California, Irvine(加州大学 Irvine分校)
;
University of California, Santa Barbara(加州大学圣巴巴拉分校)
Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events
直击核心:一种无需训练的多模态摘要方法 via 事件链
Xiaoxing You, Qiang Huang, Lingyu Li, Xiaojun Chang, Jun Yu
机构
*
School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机科学学院)
;
School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院)
;
School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)
Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
Vid-LLM:一种基于视频的紧凑型3D多模态大语言模型,具有重建-推理协同效应
Haijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu, Jiaxuan Lin, Jingrong Wang
机构
*
School of Geodesy and Geomatics, Wuhan University(武汉大学测绘学院)
;
Hubei Luojia Laboratory(湖北珞珈实验室)
;
School of Architecture and Urban Planning, Shenzhen University(深圳大学建筑与城市规划学院)
MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering
MTBench: 一种多模态时间序列基准,用于时间推理和问答
Jialin Chen, Aosong Feng, Ziyu Zhao, Juan Garza, Gaukhar Nurbek, Cheng Qin, Ali Maatouk, Leandros Tassiulas, Yifeng Gao, Rex Ying
机构
*
Yale University, New Haven, CT, USA(耶鲁大学)
;
McGill University, Montreal, QC, Canada(麦吉尔大学)
;
University of Texas Rio Grande Valley, TX, USA(德克萨斯大学里奥格兰德谷分校)
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
EDVD-LLaMA: 通过多模态大语言模型推理进行可解释的深度伪造视频检测
Haoran Sun, Chen Cai, Huiping Zhuang, Kong Aik Lee, Lap-Pui Chau, Yi Wang
机构
*
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子与电气工程系,香港理工大学)
;
School of Electrical and Electronic Engineering, Nanyang Technological University(电子与电气工程学院,南洋理工大学)
;
Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(智能工程学院,华南理工大学)
TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph
Yaru Chen, Faegheh Sardari, Peiliang Zhang, Ruohao Guo, Yang Xiang, Zhenbo Li, Wenwu Wang
机构
*
Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(视觉、语音和信号处理中心(CVSSP),萨里大学)
;
School of Computer Science and Artificial Intelligence, Wuhan University of Technology(计算机科学与人工智能学院,武汉理工大学)
;
National Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室,北京大学智能科学与技术学院)
;
College of Information and Electrical Engineering, China Agricultural University(信息与电子工程学院,中国农业大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
KU Leuven(根特大学)
;
École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)
;
Carleton University(卡尔顿大学)
E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection
Ahmad Mousavi, Yeganeh Abdollahinejad, Roberto Corizzo, Nathalie Japkowicz, Zois Boukouvalas
机构
*
Department of Mathematics and Statistics, American University, Washington, DC, USA(数学与统计学系,美国大学,华盛顿特区,美国)
;
Department of Computer Science and Mathematics, Pennsylvania State University, Harrisburg, PA, USA(计算机科学与数学系,宾夕法尼亚州立大学,哈里斯堡,宾夕法尼亚州,美国)
;
Department of Computer Science, American University, Washington, DC, USA(计算机科学系,美国大学,华盛顿特区,美国)