arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4772 篇

2601.01322 2026-01-06 cs.CV cs.AI cs.LG cs.MM eess.IV 82%

LinMU: Multimodal Understanding Made Linear

LinMU: 使多模态理解线性化

Hongjie Wang, Niraj K. Jha

机构 * Princeton University(普林斯顿大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 LinMU通过线性复杂度设计实现多模态理解,无需二次注意力模块,提升视频处理效率。

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17574 2025-12-22 cs.DC cs.LG 82%

Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing

通过GPU内部调度和资源共享实现解耦的多阶段MLLM推理

Lingxiao Zhao, Haoran Zhou, Yuezhi Che, Dazhao Cheng

机构 * Wuhan University(武汉大学)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract)

AI总结 通过GPU内部调度和资源共享实现解耦的多阶段MLLM推理,提升吞吐量和延迟性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21908 2025-12-05 cs.LG 82%

Multi-Modal Machine Learning for Early Trust Prediction in Human-AI Interaction Using Face Image and GSR Bio Signals

多模态机器学习在人机交互中早期信任预测中的应用:利用面部图像和GSR生物信号

Hamid Shamszare, Avishek Choudhury

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract)

AI总结 本研究提出多模态机器学习框架,结合面部图像和GSR生物信号,用于预测人机交互中AI或人类推荐的早期信任,通过多模态堆叠集成提升预测性能。

Comments This version contains errors in content presentation and arrangement, so it is being withdrawn until a corrected version is generated

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10068 2025-12-01 cs.CV cs.AI cs.CL 82%

Mavors: Multi-granularity Video Representation for Multimodal Large Language Model

Mavors:多粒度视频表示用于多模态大语言模型

Yang Shi, Jiaheng Liu, Yushuo Guan, Zhenhua Wu, Yuanxing Zhang, Zihao Wang, Weihong Lin, Jingyun Hua, Zekun Wang, Xinlong Chen, Bohan Zeng, Wentao Zhang, Fuzheng Zhang, Wenjing Yang, Di Zhang

机构 * Peking University(北京大学) Kling Team(Kling团队) Nanjing University(南京大学) CASIA(中国科学院自动化研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Mavors通过多粒度视频表示方法,提升多模态大语言模型在长视频理解中的时空模式保留与计算效率平衡能力。

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03179 2025-11-20 cs.CV cs.MM cs.SD eess.AS 82%

UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization

Tiantian Geng, Teng Wang, Jinming Duan, Yanfu Zhang, Weili Guan, Feng Zheng, Ling shao

机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学) School of Computer Science, University of Birmingham(计算机科学学院,伯明翰大学) Department of Computer Science, University of Hong Kong(计算机科学系,香港大学) Division of Informatics, Imaging and Data Sciences, University of Manchester(信息学、成像与数据科学系,曼彻斯特大学) William and Mary(威廉与玛丽学院) Harbin Institute of Technology(哈尔滨工业大学) UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学-Terminus AI实验室)

专题命中 视频多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Published on IEEE TPAMI

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 10280-10294, August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04351 2025-11-18 eess.SP 82%

RCMCL: A Unified Contrastive Learning Framework for Robust Multi-Modal (RGB-D, Skeleton, Point Cloud) Action Understanding

Hasan Akgul, Mari Eplik, Javier Rojas, Akira Yamamoto, Rajesh Kumar, Maya Singh

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

Comments 11 pages, 6 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17637 2025-10-29 cs.LG 82%

Causal Spatio-Temporal Prediction: An Effective and Efficient Multi-Modal Approach

Yuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen, Yunjun Gao

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21445 2025-10-27 cs.CL cs.AI cs.CV cs.LG 82%

REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring

Thanh Cong Ho, Farah Kharrat, Abderrazek Abid, Fakhri Karray

机构 * 2 Department of Electrical Computer Engineering University of Waterloo, Waterloo, ON, Canada N2L 3G1 Email 3 College of Computer Information Sciences Prince Sultan University Email

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref 2024 IEEE International Symposium on Medical Measurements and Applications (MeMeA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14203 2025-10-17 cs.CV cs.CL cs.MM 82%

Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition

Ryo Masumura, Shota Orihashi, Mana Ihori, Tomohiro Tanaka, Naoki Makishima, Taiga Yamane, Naotaka Kawata, Satoshi Suzuki, Taichi Katayama

机构 * NTT, Inc.(日本NTT公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at APSIPA ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12164 2025-10-09 cs.CV cs.AI cs.MM 82%

Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward

Yolo Yunlong Tang, Siting Xu, Teng Wang, Qin Lin, Qinglin Lu, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) Tencent Inc.(腾讯公司)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05829 2025-10-08 cs.SD cs.CV cs.LG cs.MM eess.AS 82%

FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders

Riccardo Fosco Gramaccioni, Christian Marinoni, Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello

机构 * Dept. Information Engineering, Electronics and Telecommunications (DIET), Sapienza University of Rome(信息工程、电子与电信系(DIET),罗马萨皮恩扎大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Acepted at IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01513 2025-10-03 cs.CV cs.AI cs.CL cs.IR 82%

From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding

Basem Rizk, Joel Walsh, Mark Core, Benjamin Nye

机构 * University Of Southern California(美国南加州大学)

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22646 2025-10-02 cs.CV cs.AI cs.CL 82%

Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs

Xingyu Fu, Siyi Liu, Yinuo Xu, Pan Lu, Guangqiuse Hu, Tianbo Yang, Taran Anantasagar, Christopher Shen, Yikai Mao, Yuanzhe Liu, Keyush Shah, Chung Un Lee, Yejin Choi, James Zou, Dan Roth, Chris Callison-Burch

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://deeptracereward.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08283 2025-09-23 cs.IR 82%

Serendipitous Recommendation with Multimodal LLM

Haoting Wang, Jianling Wang, Hao Li, Fangjun Yi, Mengyu Fu, Youwei Zhang, Yifan Liu, Liang Liu, Minmin Chen, Ed H. Chi, Lichan Hong, Haokai Lu

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract)

Comments Accepted by 2025 Recsys EARL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04254 2025-09-05 cs.HC 82%

MuMTAffect: A Multimodal Multitask Affective Framework for Personality and Emotion Recognition from Physiological Signals

Meisam Jamshidi Seikavandi, Fabricio Batista Narcizo, Ted Vucurevich, Andrew Burke Dittberner, Paolo Burelli

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11092 2025-08-18 cs.LG 82%

Predictive Multimodal Modeling of Diagnoses and Treatments in EHR

Cindy Shih-Ting Huang, Clarence Boon Liang Ng, Marek Rei

机构 * Imperial College London(帝国理工学院伦敦分校)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08882 2025-08-08 cs.MM cs.AI cs.CV 82%

A Novel Multimodal System to Predict Agitation in People with Dementia Within Clinical Settings: A Proof of Concept

Abeer Badawi, Somayya Elmoghazy, Samira Choudhury, Sara Elgazzar, Khalid Elgazzar, Amer Burhan

机构 * IoT Research Laboratory, Ontario Tech University(Ontario Tech 大学物联网研究实验室) Ontario Shores Centre for Mental Health Sciences(Ontario Shores 精神健康科学中心) Temerty Faculty of Medicine, University of Toronto(多伦多大学Temerty医学学院) Faculty of Science, Ontario Tech University(Ontario Tech 大学科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07261 2025-07-11 cs.LG eess.SP 82%

Robust Multimodal Learning Framework For Intake Gesture Detection Using Contactless Radar and Wearable IMU Sensors

Chunzhuo Wang, Hans Hallez, Bart Vanrumste

机构 * e-Media Research Lab(e-Media研究实验室) ESAT-STADIUS Division, KU Leuven(ESAT-STADIUS部门,KU莱顿大学) M-Group, DistriNet, Department of Computer Science, KU Leuven(M组、DistriNet、计算机科学系,KU莱顿大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments This manuscript has been submitted to a peer-reviewed journal and is currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07184 2025-06-10 cs.AI cs.CL cs.CV 82%

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images

Liangliang You, Junchi Yao, Shu Yang, Guimin Hu, Lijie Hu, Di Wang

机构 * Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学) University of Electronic Science and Technology of China(中国电子科学技术大学) University of Copenhagen(哥本哈根大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20811 2025-06-10 cs.CV cs.CL cs.MM 82%

HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models

Xiao Wang, Jingyun Hua, Weihong Lin, Yuanxing Zhang, Fuzheng Zhang, Jianlong Wu, Di Zhang, Liqiang Nie

机构 * Kuaishou Technology(快手科技) Shandong University(山东省大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02430 2025-06-05 cs.CL cs.AI cs.CV cs.LG 82%

Generative Emotion Cause Explanation in Multimodal Conversations

Lin Wang, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang, Zhitao Zhang

机构 * Northeastern University(东北大学) Shenyang Women’s and Children’s Hospital(沈阳市妇女儿童医院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01077 2025-06-03 cs.GR cs.HC 82%

TRiMM: Transformer-Based Rich Motion Matching for Real-Time multi-modal Interaction in Digital Humans

Yueqian Guo, Tianzhao Li, Xin Lyu, Jiehaolin Chen, Zhaohan Wang, Sirui Xiao, Yurun Chen, Yezi He, Helin Li, Fan Zhang

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

Comments 24 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01932 2025-05-29 cs.RO cs.LG 82%

Bridging Language, Vision and Action: Multimodal VAEs in Robotic Manipulation Tasks

Gabriela Sejnova, Michal Vavrecka, Karla Stepanova

机构 * Czech Institute of Informatics, Robotics and Cybernetics(捷克信息学、机器人学与自动控制研究所) Czech Technical University in Prague(布拉格捷克技术大学)

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract)

Comments 7 pages, 5 figures, 2 tables, conference

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14535 2025-05-21 cs.LG cs.HC 82%

Spiking Neural Networks with Temporal Attention-Guided Adaptive Fusion for imbalanced Multi-modal Learning

Jiangrong Shen, Yulin Xie, Qi Xu, Gang Pan, Huajin Tang, Badong Chen

机构 * Faculty of Electronic and Information Engineering, Xi’an Jiaotong University(电子与信息工程学院,西安交通大学) Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人研究院,西安交通大学) State Key Lab of Brain-Machine Intelligence, Zhejiang University(脑机智能国家重点实验室,浙江大学) School of Computer Science, Dalian University of Technology(计算机科学学院,大连理工大学) National Key Lab of Human-Machine Hybrid Augmented Intelligence, Xi’an Jiaotong University(人机混合增强智能国家实验室,西安交通大学)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12051 2025-05-20 cs.MM cs.AI cs.CV 82%

Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion

Yinghui Zhang, Tailin Chen, Yuchen Zhang, Zeyu Fu

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系) Institute for Analytics and Data Science, University of Essex(埃塞克斯大学分析与数据科学研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ICDMW 2024, Github: https://github.com/EvelynZ10/cmfusion

Journal ref 2024 IEEE International Conference on Data Mining Workshops (ICDMW), Abu Dhabi, United Arab Emirates, 2024, pp. 183-190

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16036 2025-03-21 cs.CV cs.AI cs.CL 82%

Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

Zhihang Liu, Chen-Wei Xie, Pandeng Li, Liming Zhao, Longxiang Tang, Yun Zheng, Chuanbin Liu, Hongtao Xie

机构 * University of Science and Technology of China(中国科学技术大学) Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室) Tsinghua University(清华大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05889 2025-03-21 cs.CV cs.AI cs.CL 82%

CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion

Shoubin Yu, Jaehong Yoon, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2025; first two authors contributed equally. Project page: https://CREMA-VideoLLM.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05092 2025-03-19 cs.CV cs.AI cs.CL 82%

Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs

Rohit Saxena, Aryo Pradipta Gema, Pasquale Minervini

机构 * University of Edinburgh(爱丁堡大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at the ICLR 2025 Workshop on Reasoning and Planning for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16755 2025-02-25 cs.CY 82%

Watch Out E-scooter Coming Through: Multimodal Sensing of Mixed Traffic Use and Conflicts Through Riders Ego-centric Views

Hiruni Nuwanthika Kegalle, Danula Hettiachchi, Jeffrey Chan, Mark Sanderson, Flora D. Salim

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract)

Comments Accepted in Proc. ACM Interactive, Mobile, Wearable and Ubiquitous Technologies,(March 2025), 23 pages. https://doi.org/10.1145/3712284

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05474 2025-01-13 cs.CL cs.AI cs.LG cs.SD eess.AS 82%

Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis

Xincheng Wang, Liejun Wang, Yinfeng Yu, Xinxin Jiao

机构 * School of Computer Science Technology , Xinjiang University, Urumqi, China Email

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments Accepted for publication by 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏