arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4772 篇

2607.16560 2026-07-21 cs.AI cs.CV cs.LG cs.MM 新提交 85%

From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence

从模态到命题:多模态智能的语言中心框架

Nadine Chang, Maying Shen, Shizhe Diao, Jialiang Wang, Jingde Chen, Thomas Breuel, Pavlo Molchanov, Rafid Mahmood, Jose M. Alvarez

机构 * NVIDIA(英伟达) University of Ottawa(渥太华大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 该研究提出多模态数据语言表示框架,将观察结果表示为原子命题,通过全局语义码本统一为共享词汇表,置于可解释空间,实现跨模态理解等,还在自动驾驶等数据上进行了展示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02945 2026-04-06 cs.DC 85%

MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference

MSAO:基于边缘-云协作的自适应模态稀疏性感知卸载框架用于高效多模态大语言模型推理

Zheming Yang, Qi Guo, Jun Wan, Jiarui Ruan, Yunqing Hu, Chang Zhao, Xiangyang Li

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract)

AI总结 本文提出MSAO框架,通过边缘-云协作和模态稀疏性分析,降低多模态大语言模型推理的延迟和资源消耗,提升吞吐量。

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14267 2026-04-06 cs.CV cs.AI cs.MM cs.SD 85%

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

DiFlowDubber:通过跨模态对齐与同步实现的离散流匹配自动化视频配音

Ngoc-Son Nguyen, Thanh V. T. Tran, Jeongsoo Choi, Hieu-Nghia Huynh-Nguyen, Truong-Son Hy, Van Nguyen

机构 * FPT Software AI Center, Vietnam(FPT软件人工智能中心,越南) KAIST, South Korea(韩国科学技术院) University of Alabama at Birmingham, USA(阿拉巴马大学伯明翰分校)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出DiFlowDubber框架,通过离散流匹配和两阶段训练策略,解决视频配音中内容准确性、表达语气、高质量音频和精确唇同步的问题。

Comments Accepted at CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28610 2026-04-01 cs.CV cs.AI cs.CL 85%

ResAdapt: Adaptive Resolution for Efficient Multimodal Reasoning

ResAdapt:面向高效多模态推理的自适应分辨率

Huanxuan Liao, Zhongtao Jiang, Yupu Hao, Yuqiao Tan, Shizhu He, Ben Wang, Jun Zhao, Kun Xu, Kang Liu

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 ResAdapt通过自适应输入分辨率框架,在保持高空间分辨率的同时提升多模态推理效率,尤其在压缩条件下显著提升性能。

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04356 2025-12-05 cs.CV cs.AI cs.CL cs.LG 85%

Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment

通过自增强对比对齐缓解多模态大语言模型中的对象和动作幻觉

Kai-Po Chang, Wei-Yuan Cheng, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国家交通大学通信工程研究所) NVIDIA

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 SANTA框架通过自增强对比对齐方法,有效缓解多模态大语言模型中的对象和动作幻觉问题。

Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026. Project page: https://kpc0810.github.io/santa/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15870 2025-10-29 cs.CV cs.AI cs.CL 85%

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang, Ligeng Zhu, Yuanhang Su, Sean Lin, An-Chieh Cheng, Zhen Wan, Jinchuan Tian, Yuming Lou, Dong Yang, Zhijian Liu, Yukang Chen, Ambrish Dantrey, Ehsan Jahangiri, Sreyan Ghosh, Daguang Xu, Ehsan Hosseini-Asl, Danial Mohseni Taheri, Vidya Murali, Sifei Liu, Yao Lu, Oluwatobi Olabiyi, Yu-Chiang Frank Wang, Rafael Valle, Bryan Catanzaro, Andrew Tao, Song Han, Jan Kautz, Hongxu Yin, Pavlo Molchanov

机构 * NVIDIA

专题命中 视频多模态 :omni-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report. Code: https://github.com/NVlabs/OmniVinci

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04651 2025-08-29 cs.IR 85%

FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation

Maolin Wang, Yutian Xiao, Binhao Wang, Sheng Zhang, Shanshan Ye, Wanyu Wang, Hongzhi Yin, Ruocheng Guo, Zenglin Xu

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)

Comments Accepted by KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17204 2025-07-24 cs.LG 85%

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation

Zixuan Wang, Jinghao Shi, Hanzhong Liang, Xiang Shen, Vera Wen, Zhiqian Chen, Yifan Wu, Zhixin Zhang, Hongyu Xiong

机构 * TikTok(字节跳动)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract)

Comments Camera Ready for ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10604 2025-06-18 cs.CV cs.AI cs.CL 85%

When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

Ruixuan Zhang, Beichen Wang, Juexiao Zhang, Zilin Bian, Chen Feng, Kaan Ozbay

机构 * New York University(纽约大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10011 2025-06-13 cs.MM cs.AI cs.CV eess.SP 85%

WDMIR: Wavelet-Driven Multimodal Intent Recognition

Weiyin Gong, Kai Zhang, Yanghai Zhang, Qi Liu, Xinjie Sun, Junyu Lu, Linbo Zhu

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(中国科学技术大学认知智能国家重点实验室) School of Computer Science, Liupanshui Normal University(黎平师范学院计算机科学学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted at IJCAI 2025, 9pages, 6figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06814 2025-05-13 cs.CV cs.AI cs.CL 85%

Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge

Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, Shoujun Zhou

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) School of Engineering, Westlake University(工程学院,西湖大学)

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12623 2025-05-05 cs.LG cs.AI cs.CV cs.MM 85%

MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network

Vrushank Ahire, Kunal Shah, Mudasir Nazir Khan, Nikhil Pakhale, Lownish Rai Sookha, M. A. Ganaie, Abhinav Dhall

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔分校) Monash University(墨尔本大学)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12231 2025-01-22 cs.CV cs.AI cs.CL 85%

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Pha Nguyen, Sailik Sengupta, Girik Malik, Arshit Gupta, Bonan Min

机构 * University of Arkansas(阿肯色大学) AWS AI Labs(亚马逊云科技AI实验室)

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19493 2024-12-30 cs.CV cs.AI cs.MM 85%

Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection

Yihao Wang, Lizhi Chen, Zhong Qian, Peifeng Li

机构 * School of Computer Science and Technology(计算机科学与技术学院) Soochow University(苏州大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18286 2024-09-30 cs.CV cs.AI cs.CL 85%

Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing

Huthaifa I. Ashqar, Ahmed Jaber, Taqwa I. Alhadidi, Mohammed Elhenawy

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00552 2024-09-04 eess.AS cs.CV cs.MM cs.SD 85%

Digit Recognition using Multimodal Spiking Neural Networks

William Bjorndahl, Jack Easton, Austin Modoff, Eric C. Larson, Joseph Camp, Prasanna Rangarajan

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments 4 pages, 2 figures, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00022 2024-09-04 cs.MM cs.AI cs.CV 85%

Detecting Misinformation in Multimedia Content through Cross-Modal Entity Consistency: A Dual Learning Approach

Zhe Fu, Kanlun Wang, Wangjiaxuan Xin, Lina Zhou, Shi Chen, Yaorong Ge, Daniel Janies, Dongsong Zhang

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to PACIS 2024. 15 pages, 3 figures

Journal ref https://aisel.aisnet.org/pacis2024/track07_secprivacy/track07_secprivacy/2

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15164 2024-01-30 cs.SD cs.CV cs.LG cs.MM eess.AS 85%

AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations

Naresh Kumar Devulapally, Sidharth Anand, Sreyasee Das Bhattacharjee, Junsong Yuan, Yu-Ping Chang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00220 2023-12-04 cs.MM cs.CL cs.CV 85%

Multi-Modal Video Topic Segmentation with Dual-Contrastive Domain Adaptation

Linzi Xing, Quan Tran, Fabian Caba, Franck Dernoncourt, Seunghyun Yoon, Zhaowen Wang, Trung Bui, Giuseppe Carenini

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at the 30th International Conference on Multimedia Modeling (MMM 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00402 2023-05-12 cs.CV cs.CL cs.MM 85%

mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Haiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li, Bin Bi, Qi Qian, Wei Wang, Guohai Xu, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou

专题命中 视频多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.MM

Journal ref ICML2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05243 2022-10-12 cs.IR 85%

Cross-modal Search Method of Technology Video based on Adversarial Learning and Feature Fusion

Xiangbin Liu, Junping Du, Meiyu Liang, Ang Li

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.05448 2018-04-17 cs.CL cs.AI cs.CV 85%

Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning

Xin Wang, Yuan-Fang Wang, William Yang Wang

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments NAACL 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21100 2025-07-30 cs.CY cs.AI cs.CV 84%

A Tactical Behaviour Recognition Framework Based on Causal Multimodal Reasoning: A Study on Covert Audio-Video Analysis Combining GAN Structure Enhancement and Phonetic Accent Modelling

Wei Meng

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments This paper introduces a structurally innovative and mathematically rigorous framework for multimodal tactical reasoning, offering a significant advance in causal inference and graph-based threat recognition under noisy conditions

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.00679 2021-08-03 cs.CV cs.MM 84%

Multimodal Feature Fusion for Video Advertisements Tagging Via Stacking Ensemble

Qingsong Zhou, Hai Liang, Zhimin Lin, Kele Xu

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments 1st place in ACM Multimedia Multimodal Video Ads Tagging Competition (2021 Tencent Advertising Algorithm Competition)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00705 2026-06-29 cs.CV 版本更新 84%

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

无训练不确定性引导的复杂视觉任务多模态大语言模型

Sanghwan Kim, Rui Xiao, Stephan Alaniz, Yongqin Xian, Zeynep Akata

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Helmholtz Munich(海德堡-慕尼黑亥姆霍兹研究中心) Google(谷歌) LTCI, Télécom Paris, Institut Polytechnique de Paris(巴黎理工大学 LTCI 实验室)

专题命中 视频多模态 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV

AI总结 提出无需训练的框架,利用MLLM内在不确定性主动引导,通过响应不确定性评分候选视觉输入,使模型自主聚焦关键信息,在视觉搜索、长视频理解等任务中达到与微调系统相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16198 2026-06-16 cs.CV 新提交 84%

GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction

GRACE: 基于接地动作中心证据增强视频多模态大语言模型用于观众情感预测

Ruoxuan Yang, Tieyuan Chen, Xiaofeng Huang, Haibing Yin, Jun Wang, Xiping Chen, Jun Yin, Xuesong Gao, Weiyao Lin

机构 * Shanghai Jiao Tong University(上海交通大学) Hangzhou Dianzi University(杭州电子科技大学) The 52nd Research Institute of China Electronics Technology Group Corporation(中国电子科技集团公司第五十二研究所) Hangzhou Bywin Technology Co., Ltd.(杭州百威科技有限公司) Zhejiang Dahua Technology Co., Ltd.(浙江大华技术股份有限公司) School of Information Science and Engineering, Shandong University(山东大学信息科学与工程学院) Haihe Laboratory of Information Technology Application Innovation(海河信息技术应用创新实验室)

专题命中 视频多模态 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV

AI总结 提出GRACE框架,通过提取时间有序的主谓宾三元组和视觉实体裁剪,增强视频MLLM对细粒度情感线索的提取与推理,在Pitts数据集上提升Qwen2.5-VL和Qwen3-VL性能。

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24845 2026-08-26 cs.CV cs.AI cs.LG 新提交 84%

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

LAION-BVD:用于多模态预训练的千万小时级开源视频数据集

Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 研究团队构建了千万小时级开源多模态视频数据集LAION-BVD,经其训练的模型在视频-文本、音频-文本及图像-文本基准上表现出色,为多模态预训练提供了大规模数据支撑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06140 2026-08-06 cs.CV cs.AI 版本更新 84%

Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion

Place-it-R1: 解锁多模态大语言模型在视频物体插入中的环境感知推理潜力

Bohai Gu, Taiyi Wu, Dazhao Du, Jian Liu, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo

机构 * HKUST(香港科技大学) Tencent Video(腾讯视频) Peking University(北京大学)

专题命中 视频多模态 :MLLM(title,abstract);分类 cs.CV、cs.AI

AI总结 Place-it-R1通过多模态大语言模型的环境感知推理能力,实现物理一致的视频物体插入,提供两种模式以平衡保真度与合理性。

Comments https://nevsnev.github.io/Place-it-R1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28375 2026-07-31 cs.AI cs.MM 新提交 84%

HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection

HyperClaim:用于视频虚假信息检测的细粒度跨模态超图推理

Xiangbo Wang, Jiasheng Zhang, Xingtong Yu, Luoqiang Lei, Delvin Ce Zhang

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM

AI总结 HyperClaim是一种用于视频虚假信息检测的判别式时间超图框架,通过细粒度跨模态超图推理,在FactGuard协议下的三个数据集上准确率优于基线方法。

Comments 13 pages, including supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25961 2026-07-30 cs.CV cs.AI 版本更新 84%

Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition

用于视频级矛盾心理和犹豫识别的交互流知识引导多模态推理

Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 研究视频级矛盾心理和犹豫识别难题,提出PRISM-AH框架,通过冻结编码器、轻量级流模型处理多模态冲突,经窗口级注释监督、知识引导大语言模型推理,在测试中取得较好宏观F1,验证推理增益可转移。

Comments 14 Pages, 1 Figure, Ambivalence/Hesitancy (AH) Video Recognition Challenge, ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏