arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2503.10200 2025-12-23 cs.CV 79%

LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents

LVAgent: 通过多轮动态协作的MLLM代理实现长视频理解

Boyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang, Yang Liu, Peng Li, Yali Wang

机构 * Shenzhen Key Lab of Computer Vision and Pattern Recognition, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳计算机视觉与模式识别重点实验室,深圳先进技术研究院,中国科学院) Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China(人工智能产业研究院(AIR),清华大学,北京,中国) Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China(计算机科学与技术系,人工智能研究院,清华大学,北京,中国) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :MLLM(title,abstract);分类 cs.CV

AI总结 LVAgent通过多轮动态协作的MLLM代理提升长视频理解性能,实现80%的准确率并在LongVideoBench上提升13.3%

Comments accepted in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16023 2025-12-19 cs.CV 79%

CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion

CoVAR: 通过多模态扩散生成视频与动作用于机器人操作

Liudi Yang, Yang Bai, George Eskandar, Fengyi Shen, Mohammad Altillawi, Dong Chen, Ziyuan Liu, Abhinav Valada

机构 * University of Freiburg(弗赖堡大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Technical University of Munich(慕尼黑技术大学) Huawei Heisenberg Research Center (Munich)(华为海森堡研究中心)

专题命中 视频多模态 :multi-modal(title);cross-modal(abstract);分类 cs.CV

AI总结 CoVAR通过多模态扩散模型生成视频与动作,解决机器人操作中动作标注不足的问题,提升视频生成质量与动作精度。

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18951 2025-12-19 cs.CV 79%

Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition

感知、对话,然后适应:面向开放世界视频识别的多模态基础模型知识迁移

Boyu Chen, Siran Chen, Kunchang Li, Qinglin Xu, Yu Qiao, Yali Wang

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) the School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出PCA框架,通过感知、对话和适应三个阶段,利用多模态知识提升开放世界视频识别的性能。

Comments 35 pages, 6 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10021 2025-12-17 cs.LG cs.AI 79%

Online Multi-modal Root Cause Identification in Microservice Systems

微服务系统中的在线多模态根本原因识别

Lecheng Zheng, Zhengzhang Chen, Haifeng Chen

机构 * University of Illinois Urbana-Champagin(伊利诺伊大学厄巴纳-香槟分校) NEC Labs America(NEC美洲实验室)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出OCEAN方法,通过多模态因果结构学习实现微服务系统中的在线根本原因识别,结合扩张卷积神经网络和图神经网络,提升因果图学习的准确性与效率。

Comments Accepted by BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13604 2025-12-16 cs.CV 79%

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

LongVie 2: 多模态可控超长视频世界模型

Jianxiong Gao, Zhaoxi Chen, Xian Liu, Junhao Zhuang, Chengming Xu, Jianfeng Feng, Yu Qiao, Yanwei Fu, Chenyang Si, Ziwei Liu

机构 * FDU(福建大学) NJU(南京大学) NTU(National University of Taiwan) NVIDIA(英伟达) THU(清华大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 LongVie 2通过多模态引导、退化感知训练和历史-上下文引导,实现了超长视频的可控生成,支持连续视频生成长达五分钟,推动视频世界建模的统一发展。

Comments Project Page: https://vchitect.github.io/LongVie2-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00752 2025-12-12 cs.CV cs.RO 79%

Multi-Modal Graph Convolutional Network with Sinusoidal Encoding for Robust Human Action Segmentation

多模态图卷积网络与正弦编码用于鲁棒的人体动作分割

Hao Xing, Kai Zhe Boey, Yuankai Wu, Darius Burschka, Gordon Cheng

机构 * Institute for Cognitive Systems(认知系统研究所) Chair of Media Technology(媒体技术教授职位) Machine Vision and Perception Group(机器视觉与感知小组) School of Computation, Information and Technology(计算、信息与技术学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出多模态图卷积网络与正弦编码方法,通过融合低帧率和高帧率数据提升人体动作分割的鲁棒性,实验表明在双手动作数据集上显著提高分割准确率。

Comments 8 pages, 5 figures, accepted in IROS25, Hangzhou, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23478 2025-12-09 cs.CV 79%

Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models

Video-R2: 通过强化一致性和基于现实的推理提升多模态语言模型

Muhammad Maaz, Hanoona Rasheed, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学) Linköping University(林肯皮大学) Australian National University(澳大利亚国立大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 Video-R2通过强化学习提升多模态模型在视频推理中的时间对齐和逻辑一致性,提高准确性和可信度。

Comments Video-R2 Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06708 2025-12-09 cs.LG cs.AI 79%

A Novel Multimodal RUL Framework for Remaining Useful Life Estimation with Layer-wise Explanations

一种新型多模态剩余寿命估计框架,结合层间解释

Waleed Razzaq, Yun-Bo Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种结合图像和时频表示的多模态RUL估计框架,通过多头注意力和多模态LRP提升模型透明度,实验证明其在不同操作条件下均优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19907 2025-12-08 cs.CV 79%

MHB: Multimodal Handshape-aware Boundary Detection for Continuous Sign Language Recognition

MHB: 多模态手形感知边界检测用于连续手语识别

Mingyu Zhao, Zhanfu Yang, Yang Zhou, Zhaoyang Xia, Can Jin, Xiaoxiao He, Dimitris N. Metaxas

机构 * Rutgers University(罗杰斯大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MHB方法,通过多模态融合结合3D骨骼特征和手形信息,提升连续手语识别的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02651 2025-12-03 cs.HC cs.CV cs.SE 79%

Real-Time Multimodal Data Collection Using Smartwatches and Its Visualization in Education

利用智能手表进行实时多模态数据采集及其在教育中的可视化

Alvaro Becerra, Pablo Villegas, Ruth Cobos

机构 * Universidad Autónoma de Madrid(马德里自治大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出利用智能手表进行实时多模态数据采集及可视化,用于教育场景中的学习分析。

Comments Accepted in Technological Ecosystems for Enhancing Multiculturality (TEEM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02609 2025-12-03 cs.RO cs.CV 79%

SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction

SAM2Grasp:通过提示条件化的时序动作预测解决多模态抓取

Shengkai Wu, Jinrong Yang, Wenqiu Luo, Linfeng Gao, Chaohui Shang, Meiyu Zhi, Mingshan Sun, Fangping Yang, Liangliang Ren, Yong Zhao

机构 * CVTE(中国中车)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 SAM2Grasp通过提示条件化的时序动作预测,解决多模态抓取中的冲突问题,实现高精度抓取性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07867 2025-12-02 cs.CV 79%

Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models

持续感知至关重要:多模态模型中时间整合失败的诊断

Zeyu Wang, Zhenzhen Weng, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CP-Bench,通过简单任务揭示多模态模型在持续感知中的时间整合缺陷,指出现有模型无法有效跨时间积累证据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22774 2025-12-01 cs.CV 79%

Alzheimer's Disease Prediction Using EffNetViTLoRA and BiLSTM with Multimodal Longitudinal MRI Data

使用EffNetViTLoRA和BiLSTM的多模态纵向MRI数据进行阿尔茨海默病预测

Mahdieh Behjat Khatooni, Mohsen Soryani

机构 * School of Computer Engineering, Iran University of Science and Technology(伊朗科学技术大学计算机工程学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出结合EffNetViTLoRA和BiLSTM的多模态模型,利用纵向MRI数据实现高精度的阿尔茨海默病预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18773 2025-12-01 cs.CV 79%

Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains

Spacewalk-18:多模态和长形式程序视频理解的基准测试,用于新领域

Zitian Tang, Rohan Myer Krishnan, Zhiqiu Yu, Chen Sun

机构 * Brown University(布朗大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 Spacewalk-18是一个用于多模态和长形式程序视频理解的新领域基准测试,通过步骤识别和视频问答任务评估模型在新领域和长时序上下文中的泛化能力。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05034 2025-11-26 cs.CV 79%

Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models

视频大模型训练后:深入探讨大型多模态模型在视频推理中的应用

Yolo Y. Tang, Jing Bi, Pinxin Liu, Zhenyu Pan, Zhangyun Tan, Qianxiang Shen, Jiani Liu, Hang Hua, Junjia Guo, Yunzhong Xiao, Chao Huang, Zhiyuan Wang, Susan Liang, Xinyi Liu, Yizhi Song, Junhua Huang, Jia-Xing Zhong, Bozheng Li, Daiqing Qi, Ziyun Zeng, Ali Vosoughi, Luchuan Song, Zeliang Zhang, Daiki Shimada, Han Liu, Jiebo Luo, Chenliang Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文深入探讨了Video-LMMs训练后处理方法,分析了监督微调、强化学习和测试时扩展等关键技术,总结了设计原则与挑战,为视频推理研究提供了统一框架。

Comments Version v1.1

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18823 2025-11-25 cs.CV 79%

VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models

VideoPerceiver: 提升视频多模态大语言模型的细粒度时序感知能力

Fufangchen Zhao, Liao Zhang, Daiqi Shi, Yuanjun Gao, Chen Ye, Yang Cai, Jian Gao, Danfeng Yan

机构 * State Key Laboratory of Networking and Switching Technology, BUPT(网络与交换技术国家重点实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 VideoPerceiver通过两阶段训练框架提升视频多模态大语言模型对细粒度动作和短暂事件的感知能力,显著优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18104 2025-11-25 cs.CV 79%

Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning

通过统一多模态伪造学习巩固扩散生成视频检测

Xiaohong Liu, Xiufeng Song, Huayu Zheng, Lei Bai, Xiaoming Liu, Guangtao Zhai

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) School of Information Science and Electronic Engineering, Shanghai Jiao Tong University(上海交通大学信息科学与电子工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Department of Computer Science and Engineering, Michigan State University(密歇根州立大学计算机科学与工程系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MM-Det++算法,通过统一多模态学习检测扩散生成视频,结合时空分支和多模态分支,提升视频伪造检测的准确性和泛化能力。

Comments Code and dataset are available at https://github.com/SparkleXFantasy/MM-Det-Plus

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17945 2025-11-25 cs.CV 79%

Test-Time Temporal Sampling for Efficient MLLM Video Understanding

测试时时间采样用于高效多模态大语言模型视频理解

Kaibin Wang, Mingbao Lin

机构 * SenseTime, China(深睿时代,中国) Rakuten, Singapore(拉结恩,新加坡)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 T3S通过测试时时间采样提高多模态大语言模型处理长视频的效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18883 2025-11-24 cs.CV 79%

Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

通用视频时间定位与生成多模态大语言模型

Zeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang, Yanfeng Wang, Weidi Xie

机构 * SAI, Shanghai Jiao Tong University(上海交通大学SAI实验室) ByteDance Seed(字节跳动种子)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出UniTime模型,利用生成多模态大语言模型实现通用视频时间定位,有效处理多类型视频并提升VideoQA任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16227 2025-11-21 cs.CV 79%

SwiTrack: Tri-State Switch for Cross-Modal Object Tracking

SwiTrack:跨模态目标跟踪的三态开关

Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) School of Mathematics, Southeast University(东南大学数学学院) Purple Mountain Laboratories(紫金山实验室)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 SwiTrack通过三态开关框架提升跨模态目标跟踪的鲁棒性和精度,实现7.2%和4.3%的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15722 2025-11-21 cs.AI 79%

Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods

多模态大语言模型中的空间推理:任务、基准和方法的调查

Weichen Liu, Qiyao Xue, Haoming Wang, Xiangyu Yin, Boyuan Yang, Wei Gao

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文调查多模态大语言模型中的空间推理问题,从认知角度分类任务和基准,分析评估方法与改进策略,揭示模型与人类推理间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14770 2025-11-20 cs.IR cs.AI 79%

ExplainRec: Towards Explainable Multi-Modal Zero-Shot Recommendation with Preference Attribution and Large Language Models

Bo Ma, LuYao Liu, ZeHua Hu, Simon Lau

机构 * Department of Software \& Microelectronics Peking University Beijing, China Department of Software \& Microelectronics Peking University Beijing, China hangli\ Department of Software \& Microelectronics Peking University Beijing, China zehua\ Department of Software \& Microelectronics Peking University Beijing, China xiaofan\ Economic Law School China University of Political Science School of Computer Science Peking University Beijing, China

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10210 2025-11-20 cs.CV 79%

MK-SGN: A Spiking Graph Convolutional Network with Multimodal Fusion and Knowledge Distillation for Skeleton-based Action Recognition

Naichuan Zheng, Hailun Xia, Zeyu Liang, Yuchen Du

机构 * Beijing Laboratory of Advanced Information Networks, Beijing Key Laboratory of Network System Architecture and Convergence, School of Information and Communication Engineering, Beijing University of Posts and Telecommunications(北京先进信息网络实验室、网络系统架构与收敛重点实验室、信息与通信工程学院、北京邮电大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14057 2025-11-19 cs.LG cs.AI 79%

A Machine Learning-Based Multimodal Framework for Wearable Sensor-Based Archery Action Recognition and Stress Estimation

Xianghe Liu, Jiajia Liu, Chuxian Xu, Minghan Wang, Hongbo Peng, Tao Sun, Jiaqi Xu

机构 * Beijing PsychTech Technology Co., Ltd.(北京心理科技技术有限公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13655 2025-11-18 cs.CV cs.LG 79%

OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation

Henry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng, Joseph Redmon, Hadrien Sablon, Ryan Park, Jacob Morrison, Alexandra Buraczynski, Karen Farley, Joshua Hansen, Andrew Howe, Patrick Alan Johnson, Mark Otterlee, Ted Schmitt, Hunter Pitelka, Stephen Daspit, Rachel Ratner, Christopher Wilhelm, Sebastian Wood, Mike Jacobi, Hannah Kerner, Evan Shelhamer, Ali Farhadi, Ranjay Krishna, Patrick Beukema

机构 * Allen Institute for AI(人工智能研究所) University of Washington(华盛顿大学) Arizona State University(亚利桑那州立大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12196 2025-11-18 cs.CV cs.HC 79%

Cross-View Cross-Modal Unsupervised Domain Adaptation for Driver Monitoring System

Aditi Bhalla, Christian Hellert, Enkelejda Kasneci

机构 * School of Social Sciences and Technology(社会科学与技术学院)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11002 2025-11-17 cs.CV 79%

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 15 pages, 12 figures. Accepted as an Oral presentation at AAAI 2026. For code and dataset, see https://zane-zyqiu.github.io/EmoVid

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10134 2025-11-14 cs.CV 79%

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

Mingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li, Qi Zeng, Yifan Zhang, Ju Xin, Rongtao Xu, Jiguang Zhang, Xiaopeng Zhang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23990 2025-11-13 cs.AI 79%

Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding

Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri, Nicholas R. Waytowich, Xiaomin Lin, Tinoosh Mohsenin

机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02182 2025-11-05 cs.CV 79%

Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models

Jinhwan Seo, Yoonki Cho, Junhyug Noh, Sung-eui Yoon

机构 * KAIST(韩国科学技术院) Ewha Womans University(成均馆大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 1st place winner of Grounded Videoqa track at the ICCV2025 Perception Test

详情

展开后加载摘要…

URL PDF HTML 收藏