arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4772 篇

2604.17422 2026-04-21 cs.CV cs.MM 86%

Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding

聚焦何处:用于长视频理解的查询调节多模态关键帧选择

Shaoguang Wang, Weiyu Guo, Ziyang Chen, Xuming Hu, Hui Xiong

机构 * Department of CSE, HKUST(香港科技大学计算机科学与工程系)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.MM

AI总结 本文提出Q-Gate框架,通过动态模态路由解决长视频理解中关键帧选择问题,有效抑制模态噪声,提升多模态大语言模型的推理能力。

Comments 9 pages, 7 figures, 9 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28696 2026-03-31 cs.CV cs.AI 86%

AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding

AdaptToken: 基于熵的自适应令牌选择用于多模态大语言模型长视频理解

Haozhe Qi, Kevin Qu, Mahdi Rad, Rui Wang, Alexander Mathis, Marc Pollefeys

机构 * Microsoft Spatial AI Lab(微软空间人工智能实验室) EPFL(瑞士联邦理工学院洛桑) ETH Zurich(苏黎世联邦理工学院)

专题命中 视频多模态 :MLLM(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 AdaptToken通过利用模型自不确定性实现全局控制信号,提升多模态大语言模型对长视频的理解能力,通过熵信号分配令牌预算并支持提前停止,提升准确率并减少推理时间。

Comments Project page: https://haozheqi.github.io/adapt-token

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03722 2025-08-07 cs.CV cs.AI 86%

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi, Feng Chen, Xinming Wang, Jun Xie

机构 * Lenovo Research(联想研究院) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Institute of Automation, CAS(中国科学院自动化研究所) Macau University of Science and Technology(澳门科学理工学院) School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07138 2025-03-03 cs.CV cs.CL cs.LG 86%

Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content

Girish A. Koushik, Diptesh Kanojia, Helen Treharne

机构 * University of Surrey(萨里大学) NICE Research(NICE研究所) Surrey Centre for Cyber Security(萨里网络安全中心)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted to the MM4SG Workshop at the WebConf 2025

Journal ref Companion Proceedings of the ACM Web Conference 2025 (WWW Companion '25), April 28-May 2, 2025, Sydney, NSW, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16552 2024-07-25 cs.CV cs.MM 86%

MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues

Liyun Zhang

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.12059 2021-02-02 cs.CV cs.CL 86%

VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs

Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, Lorenzo Torresani

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10212 2025-11-14 cs.CV 86%

Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization

Ashutosh Anshul, Shreyas Gopal, Deepu Rajan, Eng Siong Chng

机构 * College of Computing and Data Science(计算与数据科学学院) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Under Review, Multimodal Deepfake detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18314 2026-02-13 q-bio.QM cs.LG q-bio.NC 86%

BrainSymphony: A parameter-efficient multimodal foundation model for brain dynamics with limited data

BrainSymphony: 一种参数高效、多模态的基础模型,用于在有限数据下的脑动态

Moein Khajehnejad, Forough Habibollahi, Devon Stoliker, Adeel Razi

机构 * Turner Institute for Brain and Mental Health(大脑与心理健康Turner研究所) School of Psychological Sciences, Monash University(墨尔本大学心理学科学学院) Cortical Labs(皮层实验室) CIFAR Azrieli Global Scholars Program(CIFAR阿兹里埃利全球学者计划)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title)

AI总结 BrainSymphony是一种参数高效、多模态的基础模型,通过整合fMRI和扩散MRI数据,实现有限数据下的脑动态分析,优于更大模型并提升神经科学应用。

Comments 32 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23373 2026-08-25 cs.AI stat.AP 新提交 85%

Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting

模态应相互协作:用于长期流感预测的双流多模态学习

Seyed Mohammad Hossein Hashemi, Mohsen Hooshmand, Parvin Razzaghi

机构 * Institute for Advanced Studies in Basic Sciences (IASBS)(基础科学高级研究所)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract,abstract_cn);分类 cs.AI

AI总结 该研究针对长期流感预测问题,提出双流注意力(DSA)多模态框架,通过双向跨模态注意力耦合数值与文本流,在多个数据集上实现优于iTransformer等基线的预测性能,且方法具有鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17583 2026-08-19 cs.CL 新提交 85%

Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study

使用多模态语言模型对TikTok上的有害内容暴露情况进行审计:一项跨国、按年龄分层的研究

Hamidreza Saffari, Francesco Pierri

机构 * Politecnico di Milano(米兰理工大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

AI总结 本研究使用多模态大语言模型Gemini 2.5 Flash,在法、意、瑞三国对TikTok开展跨国年龄分层审计,发现关键词搜索会大幅提升有害内容占比,意国各年龄组有害内容占比最高,平台安全过滤器低估了明确有害内容。

Comments 20 pages, 16 figures, 14 tables. Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04824 2026-08-07 cs.CV 85%

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

SOVABench:多模态大语言模型的车辆监控动作检索基准

Oriol Rabasseda, Zenjie Li, Kamal Nasrollahi, Sergio Escalera

机构 * Milestone Systems A/S(Milestone Systems公司) Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心) Aalborg Universitet(奥胡斯大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 SOVABench为多模态大语言模型提供车辆监控动作检索基准,通过定义两种评估协议评估跨动作区分和时间方向理解,展示了模型在复杂监控任务中的性能。

Comments This work has been accepted at Real World Surveillance: Applications and Challenges, 6th (in WACV Workshops)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16978 2026-06-17 cs.CV 版本更新 85%

A Benchmark for Omni-Modal Reasoning in Long Videos

长视频全模态推理基准

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou, Sahal Shaji Mullappilly, Mohammad Almansoori, Noor Ahsan, Beknur Kalmakhanbet, Sambal Shikhar, Rishabh Lalla, Jean Lahoud, Mariette Awad, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) American University of Beirut(贝鲁特美国大学) Linköping University(利尔贝里大学)

专题命中 视频多模态 :omni-modal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CV

AI总结 提出LongShOTBench基准,用于评估长视频中视觉、语音和环境音频的全模态推理,并引入无训练的全模态证据搜索代理LongShOTAgent,在105个模型上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03920 2026-06-03 cs.CV 85%

Benchmarking Visual State Tracking in Multimodal Video Understanding

多模态视频理解中的视觉状态追踪基准测试

Sihyun Yu, Nanye Ma, Pinzhi Huang, Hyunseok Lee, Shusheng Yang, June Suk Choi, Ellis Brown, Oscar Michel, Boyang Zheng, Jinwoo Shin, Saining Xie

机构 * New York University(纽约大学) KAIST(韩国科学技术院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 提出VSTAT基准,通过需要连续感知和整合整个视频流的问题评估多模态大语言模型的视觉状态追踪能力,发现当前模型远低于人类表现,失败主要源于视觉感知而非文本推理。

Comments Website: https://vision-x-nyu.github.io/vstat-site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11283 2026-06-02 cs.CV 85%

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

多模态大语言模型驱动的视频翻译:面向角色的综述

Bingzheng Qu, Kehai Chen, Xuefeng Bai, Min Zhang

机构 * School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文通过面向角色的分类法,系统综述了多模态大语言模型在视频翻译中的应用,将其分为语义推理器、表达执行器和视觉合成器三个功能角色,并总结了数据集、基准和评估指标,指出了端到端视频翻译的挑战与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22185 2026-05-22 cs.CV cs.LG 85%

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

增强多模态大语言模型以用于安全关键驾驶视频分析

Tomaso Trinci, Henrique Piñeiro Monteagudo, Leonardo Taccari

机构 * Verizon Connect

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本研究通过融合降采样视频帧与同步高频 telemetry 数据及专用计算机视觉模型的语义信息,提升多模态大语言模型在安全关键驾驶场景中的感知与推理能力,从而更准确地识别和描述现实驾驶中的安全关键事件。

Comments Accepted at the 2026 IEEE International Conference on Intelligent Transportation Systems (ITSC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18431 2026-05-20 cs.CV 85%

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

协同视见:基于多模态大语言模型的多机器人协作自体空间推理

Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Hao Shi, Yi Zhou, M. Saquib Sarfraz, Danda Pani Paudel, Luc Van Gool

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Hunan University(湖南大学) University of Oxford(牛津大学) Zhejiang University(浙江大学) ETH Zurich(苏黎世联邦理工学院) Ant Group(蚂蚁集团)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文研究了多机器人协作动态空间推理问题,提出了首个针对该任务的基准CoopSR以及多机器人自体问答数据集EgoTeam,通过引入SP-CoR框架实现了细粒度的协作空间推理,显著提升了多机器人协作推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22492 2026-04-27 eess.IV cs.CV 85%

MTT-Bench: Predicting Social Dominance in Mice via Multimodal Large Language Models

MTT-Bench:通过多模态大语言模型预测小鼠的社会支配

Yunquan Chen, Haoyu Chen

机构 * Department of Communication Systems, KTH Royal Institute of Technology(通信系统系,皇家理工学院) CMVS, University of Oulu(奥卢大学CMVS)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出MTT-Bench基准,利用多模态大语言模型分析小鼠行为视频,预测社会支配等级,无需显式标签进行零样本推理,展示了在乙学和社会行为分析中的应用潜力。

Comments 8 pages, 2 figures. Submitted to conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12961 2025-02-28 cs.CV 85%

Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Zuyan Liu, Yuhao Dong, Ziwei Liu, Winston Hu, Jiwen Lu, Yongming Rao

机构 * Tsinghua University(清华大学) Tencent(腾讯)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00832 2024-12-03 cs.CV 85%

EventGPT: Event Stream Understanding with Multimodal Large Language Models

Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang, Xin Meng, Fei Richard Yu, Xiangyang Ji, Ming Li

机构 * Xidian University(西安电子科技大学) Tsinghua University(清华大学) UCAS(中国科学院大学) Peking University(北京大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy(SZ)(广东省人工智能与数字经济实验室(深圳))

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12828 2024-10-18 cs.CV cs.LG 85%

GCM-Net: Graph-enhanced Cross-Modal Infusion with a Metaheuristic-Driven Network for Video Sentiment and Emotion Analysis

Prasad Chaudhari, Aman Kumar, Chandravardhan Singh Raghaw, Mohammad Zia Ur Rehman, Nagendra Kumar

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17880 2024-06-27 cs.CV 85%

MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval

Weitong Cai, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu

专题命中 视频多模态 :MLLM(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04154 2024-01-10 cs.CV cs.AI cs.LG cs.MM cs.SD eess.AS 85%

Efficient Selective Audio Masked Multimodal Bottleneck Transformer for Audio-Video Classification

Wentao Zhu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by WACV 2024; well-formatted PDF is in https://drive.google.com/file/d/1qvW52lamsvNGMCqPS7q8g8L4NaR_LlbR/view?usp=sharing. arXiv admin note: text overlap with arXiv:2401.04023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17642 2023-10-27 cs.RO cs.CV cs.LG 85%

Drive Anywhere: Generalizable End-to-end Autonomous Driving with Multi-modal Foundation Models

Tsun-Hsuan Wang, Alaa Maalouf, Wei Xiao, Yutong Ban, Alexander Amini, Guy Rosman, Sertac Karaman, Daniela Rus

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Project webpage: https://drive-anywhere.github.io Explainer video: https://www.youtube.com/watch?v=4n-DJf8vXxo&feature=youtu.be

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07646 2022-07-18 cs.CV cs.LG 85%

Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Rui Qian, Yeqing Li, Zheng Xu, Ming-Hsuan Yang, Serge Belongie, Yin Cui

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12485 2022-03-30 cs.CV 85%

CroMo: Cross-Modal Learning for Monocular Depth Estimation

Yannick Verdié, Jifei Song, Barnabé Mas, Benjamin Busam, Aleš Leonardis, Steven McDonagh

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted for publication at CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07515 2021-12-15 cs.CV cs.AI cs.CL cs.MM 85%

CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising

Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao, Hongyang Chao, Tao Mei

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACM Multimedia 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10949 2021-10-22 cs.CV 85%

Multimodal Learning using Optimal Transport for Sarcasm and Humor Detection

Shraman Pramanick, Aniket Roy, Vishal M. Patel

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to WACV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04124 2021-09-23 cs.CV 85%

Parameter Efficient Multimodal Transformers for Video Representation Learning

Sangho Lee, Youngjae Yu, Gunhee Kim, Thomas Breuel, Jan Kautz, Yale Song

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Accepted to ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.03152 2017-04-12 cs.CV 85%

Deep Multimodal Representation Learning from Temporal Data

Xitong Yang, Palghat Ramesh, Radha Chitta, Sriganesh Madhvanath, Edgar A. Bernal, Jiebo Luo

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments To appear in CVPR 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10024 2026-07-20 cs.CV cs.AI cs.LG 版本更新 85%

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

LVSum:一个用于时间感知长视频摘要的基准测试

Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan

机构 * Apple(苹果公司)

专题命中 视频多模态 :MLLM(summary_cn,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出LVSum基准测试,用于评估长视频摘要中时间对齐的性能,通过引入新的评估指标揭示现有MLLM在时间理解上的系统性差距。

Comments 25 pages, 5 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏