arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2509.10802 2025-09-16 q-fin.RM cs.CL cs.LG q-fin.CP 79%

Why Bonds Fail Differently? Explainable Multimodal Learning for Multi-Class Default Prediction

Yi Lu, Aifan Ling, Chaoqun Wang, Yaxin Xu

机构 * School of Economics and Finance, Shanghai International Studies University(经济金融学院,上海国际问题研究大学) School of AI and Advanced Computing, Xi’an Jiaotong-Liverpool University(人工智能与先进计算学院,西安交通大学利物浦大学) School of Foreign Studies, Shanghai University of Finance and Economics(外国语言学院,上海金融学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14581 2025-09-11 cs.MM eess.IV 79%

Memory-Anchored Multimodal Reasoning for Explainable Video Forensics

Chen Chen, Runze Li, Zejun Zhang, Pukun Zhao, Fanqing Zhou, Longxiang Wang, Haojian Huang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06831 2025-09-09 cs.CV 79%

Leveraging Generic Foundation Models for Multimodal Surgical Data Analysis

Simon Pezold, Jérôme A. Kurylec, Jan S. Liechti, Beat P. Müller, Joël L. Lavanchy

机构 * Department of Biomedical Engineering, University of Basel, Allschwil, Switzerland(巴塞尔大学生物医学工程系) Clarunis – University Digestive Health Care Center Basel, Basel, Switzerland(Clarunis – 巴塞尔大学消化健康医疗中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages, 3 figures; accepted at ML-CDS @ MICCAI 2025, Daejeon, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06422 2025-09-09 cs.CV 79%

Phantom-Insight: Adaptive Multi-cue Fusion for Video Camouflaged Object Detection with Multimodal LLM

Hua Zhang, Changjiang Luo, Ruoyu Chen

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 视频多模态 :multimodal(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06389 2025-09-09 cs.SD cs.AI 79%

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

Xiaoran Yang, Jianxuan Yang, Xinyue Guo, Haoyu Wang, Ningning Pan, Gongping Huang

机构 * School of Electronic Information, Wuhan University, Wuhan, China(武汉大学电子信息学院) MiLM Plus, Xiaomi Inc., Wuhan, China(小米公司MiLM Plus团队)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05265 2025-09-08 cs.AI 79%

MMoE: Robust Spoiler Detection with Multi-modal Information and Domain-aware Mixture-of-Experts

Zinan Zeng, Sen Ye, Zijian Cai, Heng Wang, Yuhan Liu, Haokai Zhang, Minnan Luo

机构 * Xi’an Jiaotong University(西安交通大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02807 2025-09-04 cs.CV 79%

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

Mennatullah Siam

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Work under review in NeurIPS 2025 with the title "Are we using Motion in Referring Segmentation? A Motion-Centric Evaluation"

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01338 2025-09-03 cs.AI 79%

Conformal Predictive Monitoring for Multi-Modal Scenarios

Francesca Cairoli, Luca Bortolussi, Jyotirmoy V. Deshmukh, Lars Lindemann, Nicola Paoletti

机构 * University of Trieste, Trieste, Italy(特里埃斯特大学) University of Southern California, Los Angeles, California(南加州大学) King's College London, London, United Kingdom(伦敦国王学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04817 2025-09-03 cs.CV 79%

LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Hongjie Zhang, Lu Dong, Yi Liu, Yifei Huang, Yali Wang, Limin Wang, Yu Qiao

机构 * OpenGVLab, Shanghai AI Laboratory, China(OpenGVLab,上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Honor Device Co.,Ltd(荣耀设备有限公司) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) Nanjing University(南京大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15825 2025-09-03 cs.CL q-fin.ST 79%

Enhancing Cryptocurrency Sentiment Analysis with Multimodal Features

Chenghao Liu, Aniket Mahanti, Ranesh Naha, Guanghao Wang, Erwann Sbai

机构 * Department of Computer Science, The University of Auckland(计算机科学系,奥克兰大学) School of Information Systems, Queensland University of Technology(信息系统学院,昆士兰技术大学) Department of Economics, The University of Auckland(经济学系,奥克兰大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01591 2025-09-03 cs.CV 79%

Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval

Adriano Fragomeni, Dima Damen, Michael Wray

机构 * School of Computer Science University of Bristol(计算机科学学院英国布里斯托尔大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18284 2025-08-27 cs.LG cs.AI cs.SY eess.SY 79%

Multi-Modal Drift Forecasting of Leeway Objects via Navier-Stokes-Guided CNN and Sequence-to-Sequence Attention-Based Models

Rahmat K. Adesunkanmi, Alexander W. Brandt, Masoud Deylami, Gustavo A. Giraldo Echeverri, Hamidreza Karbasian, Adel Alaeddini

机构 * Departments of Mechanical and Electrical Engineering, Southern Methodist University(机械与电气工程系,南方 Methodist 大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17890 2025-08-26 cs.CV 79%

UniAPO: Unified Multimodal Automated Prompt Optimization

Qipeng Zhu, Yanzhe Chen, Huasong Zhong, Yan Li, Jie Chen, Zhixin Zhang, Junping Zhang, Zhenheng Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17817 2025-08-26 cs.CV 79%

TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration

Meiqi Gong, Hao Zhang, Xunpeng Yi, Linfeng Tang, Jiayi Ma

机构 * Electronic Information School, Wuhan University, Wuhan 430072, China(武汉大学电子信息学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14395 2025-08-21 cs.HC cs.AI 79%

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding

Running Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang, Handi Chen, Weipeng Deng, Luyao Jin, Xiaojuan Qi, Xun Qian, Edith C. H. Ngai

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Google(谷歌)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to UIST 2025. Project website: https://zhaorunning.github.io/NoteIt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13692 2025-08-20 cs.CV 79%

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

Keliang Li, Hongze Shen, Hao Shi, Ruibing Hou, Hong Chang, Jie Huang, Chenghao Jia, Wen Wang, Yiling Wu, Dongmei Jiang, Shiguang Shan, Xilin Chen

机构 * Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, China(中国科学院智能信息处理重点实验室(中国科学院)、计算技术研究所、中国)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10922 2025-08-18 cs.CV 79%

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 20 pages,6 figures,survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10828 2025-08-15 cs.RO cs.AI 79%

A Multimodal Neural Network for Recognizing Subjective Self-Disclosure Towards Social Robots

Henry Powell, Guy Laban, Emily S. Cross

机构 * Amazon(亚马逊) School of Psychology and Neuroscience, University of Glasgow(心理学与神经科学学院,格拉斯哥大学) Ben-Gurion University of the Negev(内盖夫本·古里安大学) University of Cambridge(剑桥大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02935 2025-08-13 cs.CL 79%

Dynamic Graph Neural ODE Network for Multi-modal Emotion Recognition in Conversation

Yuntao Shou, Tao Meng, Wei Ai, Keqin Li

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学教育部长江网络与网络安全重点实验室) College of Computer and Mathematics, Central South University of Forestry and Technology(中南林业科技大学计算机与数学学院) Department of Computer Science, State University of New York(纽约州立大学新帕尔茨分校计算机科学系)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06939 2025-08-12 cs.AI cs.LG 79%

Intrinsic Explainability of Multimodal Learning for Crop Yield Prediction

Hiba Najjar, Deepak Pathak, Marlon Nuske, Andreas Dengel

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03694 2025-08-06 cs.CV 79%

LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

Jianxiong Gao, Zhaoxi Chen, Xian Liu, Jianfeng Feng, Chenyang Si, Yanwei Fu, Yu Qiao, Ziwei Liu

机构 * Nanjing University(南京大学) Fudan University(复旦大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室) NVIDIA(NVIDIA公司) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Project page: https://vchitect.github.io/LongVie-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02179 2025-08-05 cs.CV 79%

Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning

Wenbo Xu, Wei Lu, Xiangyang Luo

机构 * School of Computer Science and Engineering, MoE Key Laboratory of Information Technology, Guangdong Province Key Laboratory of Information Security Technology, Sun Yat-sen University(计算机科学与工程学院、信息科技教育部重点实验室、广东省信息安全技术重点实验室、中山大学) State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages,4 figures. arXiv admin note: text overlap with arXiv:2507.16596

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01701 2025-08-05 cs.LG cs.AI 79%

MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model

Asmit Bandyopadhyay, Rohit Basu, Tanmay Sen, Swagatam Das

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01168 2025-08-05 cs.MM 79%

Graph-based Interaction Augmentation Network for Robust Multimodal Sentiment Analysis

Hu Zhangfeng, Shi mengxin

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06603 2025-08-04 cs.CV 79%

Cross-Modal Dual-Causal Learning for Long-Term Action Recognition

Xu Shaowu, Jia Xibin, Gao Junyu, Sun Qianmei, Chang Jing, Fan Chao

机构 * Beijing University of Technology(北京理工大学) Chinese Academy of Sciences(中国科学院) Capital Medical University(首都医科大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22878 2025-07-31 cs.IR cs.CL cs.CY 79%

GeoOutageKG: A Multimodal Geospatiotemporal Knowledge Graph for Multiresolution Power Outage Analysis

Ethan Frakes, Yinghui Wu, Roger H. French, Mengjie Li

机构 * University of Central Florida(佛罗里达中央大学) Case Western Reserve University(凯斯西储大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to the 24th International Semantic Web Conference Resource Track (ISWC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21649 2025-07-30 cs.CV 79%

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

Shibo Gao, Peipei Yang, Haiyang Guo, Yangyang Liu, Yi Chen, Shuai Li, Han Zhu, Jian Xu, Xu-Yao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 视频多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10541 2025-07-30 cs.IR cs.AI 79%

Multi-Modal Hypergraph Enhanced LLM Learning for Recommendation

Xu Guo, Tong Zhang, Yuanzhi Wang, Chenxu Wang, Fuyun Wang, Xudong Wang, Xiaoya Zhang, Xin Liu, Zhen Cui

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) Shituoyun (Nanjing) Technology Co., Ltd(石图云(南京)科技有限公司) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 12 pages, 4 figures, submitted to IEEE Transactions on Knowledge and Data Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20451 2025-07-29 cs.AI 79%

STARN-GAT: A Multi-Modal Spatio-Temporal Graph Attention Network for Accident Severity Prediction

Pritom Ray Nobin, Imran Ahammad Rifat

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20163 2025-07-29 cs.CV 79%

Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning

Zeyu Xi, Haoying Sun, Yaofei Wu, Junchi Yan, Haoran Zhang, Lifang Wu, Liang Wang, Changwen Chen

机构 * Beijing University of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学) Chinese Academy of Sciences(中国科学院) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025 (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏