arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2507.18104 2025-07-28 cs.CV q-bio.NC 79%

A Multimodal Seq2Seq Transformer for Predicting Brain Responses to Naturalistic Stimuli

Qianyi He, Yuan Chang Leong

机构 * Data Science Institute University of Chicago(芝加哥大学数据科学研究所) Department of Psychology, Neuroscience Institute University of Chicago(芝加哥大学心理学系、神经科学研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17050 2025-07-24 cs.CV 79%

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

Tz-Ying Wu, Tahani Trigui, Sharath Nittur Sridhar, Anand Bodas, Subarna Tripathi

机构 * Intel(英特尔)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to CVAM Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13482 2025-07-21 cs.LG cs.CV 79%

Improving Out-of-distribution Human Activity Recognition via IMU-Video Cross-modal Representation Learning

Seyyed Saeid Cheshmi, Buyao Lyu, Thomas Lisko, Rajesh Rajamani, Robert A. McGovern, Yogatheesan Varatharajah

机构 * Department of Computer Science & Engineering University of Minnesota(计算机科学与工程系明尼苏达大学) Department of Mechanical Engineering University of Minnesota(机械工程系明尼苏达大学) Department of Neurosurgery University of Minnesota(神经外科系明尼苏达大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13403 2025-07-21 cs.CV cs.LG 79%

UL-DD: A Multimodal Drowsiness Dataset Using Video, Biometric Signals, and Behavioral Data

Morteza Bodaghi, Majid Hosseini, Raju Gottumukkala, Ravi Teja Bhupatiraju, Iftikhar Ahmad, Moncef Gabbouj

机构 * University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校) Tietoevry(蒂奥维瑞) Tampere University(塔尔库大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12666 2025-07-18 cs.AI cs.LG 79%

Fly, Fail, Fix: Iterative Game Repair with Reinforcement Learning and Large Multimodal Models

Alex Zook, Josef Spjut, Jonathan Tremblay

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published at Reinforcement Learning and Video Games workshop https://sites.google.com/view/rlvg-workshop-2025/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15681 2025-07-18 cs.CV 79%

Vidi: Large Multimodal Models for Video Understanding and Editing

Vidi Team, Celong Liu, Chia-Wen Kuo, Dawei Du, Fan Chen, Guang Chen, Jiamin Yuan, Lingxi Zhang, Lu Guo, Lusha Li, Longyin Wen, Qingyu Chen, Rachel Deng, Sijie Zhu, Stuart Siew, Tong Jin, Wei Lu, Wen Zhong, Xiaohui Shen, Xin Gu, Xing Mei, Xueqiong Qu, Zhenfang Chen

机构 * Intelligent Editing Team(智能编辑团队) Intelligent Creation, ByteDance Inc.(智能创作,字节跳动公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09617 2025-07-15 cs.AI cs.RO 79%

Bridging Bots: from Perception to Action via Multimodal-LMs and Knowledge Graphs

Margherita Martorana, Francesca Urgese, Mark Adamik, Ilaria Tiddi

机构 * Vrije Universiteit Amsterdam(范·艾克大学阿姆斯特丹)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07938 2025-07-11 cs.MM 79%

Multimodal Framework for Explainable Autonomous Driving: Integrating Video, Sensor, and Textual Data for Enhanced Decision-Making and Transparency

Abolfazl Zarghani, Amirhossein Ebrahimi, Amir Malekesfandiari

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06072 2025-07-09 cs.CV 79%

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding

Tongtong Cheng, Rongzhen Li, Yixin Xiong, Tao Zhang, Jing Wang, Kai Liu

机构 * Department of Computer Science, Chongqing University, China(重庆大学计算机科学系) National Elite Institute of Engineering, Chongqing University, China(重庆大学工程精英研究院) College of Computer Science and Technology, National University of Deffense Technology, China(国防科技大学计算机科学与技术学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05899 2025-07-09 cs.CV 79%

What You Have is What You Track: Adaptive and Robust Multimodal Tracking

Yuedong Tan, Jiawei Shao, Eduard Zamfir, Ruanjun Li, Zhaochong An, Chao Ma, Danda Paudel, Luc Van Gool, Radu Timofte, Zongwei Wu

机构 * TeleAI, China Telecom(TeleAI,中国电信) Computer Vision Lab, CAIDAS & IFI, University of Wurzburg(计算机视觉实验室,CAIDAS与IFI,乌尔姆大学) INSAIT, Sofia University(INSAIT,索菲亚大学) ShanghaiTech University(上海科技大学) University of Copenhagen(哥本哈根大学) AI Institute, Shanghai Jiao Tong University(人工智能研究院,上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICCV2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02626 2025-07-04 cs.MM 79%

VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning

Siran Chen, Boyu Chen, Chenyun Yu, Yuxiao Luo, Ouyang Yi, Lei Cheng, Chengxiang Zhuo, Zang Li, Yali Wang

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14171 2025-07-04 cs.CV 79%

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han, Li Fei-Fei, Saining Xie

机构 * New York University(纽约大学) Yale University(耶鲁大学) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://vision-x-nyu.github.io/thinking-in-space.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17457 2025-06-24 cs.CV 79%

When Every Millisecond Counts: Real-Time Anomaly Detection via the Multimodal Asynchronous Hybrid Network

Dong Xiao, Guangyao Chen, Peixi Peng, Yangru Huang, Yifan Zhao, Yongxing Dai, Yonghong Tian

机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学) Department of Software and Microelectronics, Peking University(软件与微电子系,北京大学) School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University(虚拟现实技术与系统国家重点实验室,北航软件学院) Peng Cheng Laboratory(鹏城实验室) Baidu Inc(百度公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICML 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16401 2025-06-23 cs.CY cs.CV 79%

TrajSceneLLM: A Multimodal Perspective on Semantic GPS Trajectory Analysis

Chunhou Ji, Qiumeng Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Under review for ACM SIGSPATIAL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12456 2025-06-23 cs.CV 79%

Demographics-Informed Neural Network for Multi-Modal Spatiotemporal forecasting of Urban Growth and Travel Patterns Using Satellite Imagery

Eugene Kofi Okrah Denteh, Andrews Danyo, Joshua Kofi Asamoah, Blessing Agyei Kyem, Armstrong Aboah

机构 * Department of Civil, Construction and Environmental Engineering, North Dakota State University(土木、建筑与环境工程系,北达科塔州立大学)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23623 2025-06-23 cs.CV 79%

On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection

Xiufeng Song, Xiao Guo, Jiache Zhang, Qirui Li, Lei Bai, Xiaoming Liu, Guangtao Zhai, Xiaohong Liu

机构 * Shanghai Jiao Tong University(上海交通大学) Michigan State University(密歇根州立大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 9 figures, published in NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00631 2025-06-17 cs.LG cs.AI 79%

TrialBench: Multi-Modal Artificial Intelligence-Ready Clinical Trial Datasets

Jintai Chen, Yaojun Hu, Mingchen Cai, Yingzhou Lu, Yue Wang, Xu Cao, Miao Lin, Hongxia Xu, Jian Wu, Cao Xiao, Jimeng Sun, Yuqiang Li, Lucas Glass, Kexin Huang, Marinka Zitnik, Tianfan Fu

机构 * AI Thrust, Information Hub, HKUST(GZ)(香港科技大学(广州)人工智能 thrust 与信息中心) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Medicine, Stanford University(斯坦福大学医学院) Computer Science Department, UIUC(伊利诺伊大学厄巴纳-香槟分校计算机科学系) Medical Big Data Center, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University(广东省人民医院(广东省医学科学院)医学大数据中心,南方医科大学) Innovation Institute for Artificial Intelligence in Medicine of Zhejiang University, College of Pharmaceutical Sciences, Zhejiang University(浙江大学人工智能医学创新研究院,浙江大学药学院) The Second Affiliated Hospital, Zhejiang University School of Medicine(浙江大学医学院附属第二医院) GE HealthCare(通用电气医疗) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) IQVIA Computer Science Department, Stanford University(斯坦福大学计算机科学系) Informatics, Harvard Medical School, Harvard University(哈佛医学院,哈佛大学informatics部门) State Key Laboratory for Novel Software Technology at Nanjing University, School of Computer Science, Nanjing University(南京大学新型软件技术国家重点实验室,南京大学计算机科学学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments accepted by Nature Scientific Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16998 2025-06-12 cs.CV 79%

Understanding Long Videos with Multimodal Language Models

Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya, Michael S. Ryoo

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages (main paper), 7 pages appendix. ICLR 2025 conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09345 2025-06-12 cs.CV 79%

An Effective End-to-End Solution for Multimodal Action Recognition

Songping Wang, Xiantao Hu, Yueming Lyu, Caifeng Shan

机构 * School of Intelligence Science and Technology, Nanjing University, China(智能科学与技术学院,南京大学) PCA-Lab, School of Computer Science and Engineering, Nanjing University of Science and Technology, China(PCA实验室,计算机科学与工程学院,南京理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18253 2025-06-10 cs.LG cs.AI q-bio.QM 79%

Multimodal Integration of Longitudinal Noninvasive Diagnostics for Survival Prediction in Immunotherapy Using Deep Learning

Melda Yeghaian, Zuhir Bodalal, Daan van den Broek, John B A G Haanen, Regina G H Beets-Tan, Stefano Trebeschi, Marcel A J van Gerven

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref Journal of the American Medical Informatics Association, 2025;, ocaf074

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06205 2025-06-09 cs.RO cs.AI 79%

Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning

Sheng Chen, Peiyu He, Jiaxin Hu, Ziyang Liu, Yansheng Wang, Tao Xu, Chi Zhang, Chongchong Zhang, Chao An, Shiyu Cai, Duo Cao, Kangping Chen, Shuai Chu, Tianwei Chu, Mingdi Dan, Min Du, Weiwei Fang, Pengyou Fu, Junkai Hu, Xiaowei Jiang, Zhaodi Jiang, Fuxuan Li, Jun Li, Minghui Li, Mingyao Li, Yanchang Li, Zhibin Li, Guangming Liu, Kairui Liu, Lihao Liu, Weizhi Liu, Xiaoshun Liu, Yufei Liu, Yunfei Liu, Qiang Lu, Yuanfei Luo, Xiang Lv, Hongying Ma, Sai Ma, Lingxian Mi, Sha Sa, Hongxiang Shu, Lei Tian, Chengzhi Wang, Jiayu Wang, Kaijie Wang, Qingyi Wang, Renwen Wang, Tao Wang, Wei Wang, Xirui Wang, Chao Wei, Xuguang Wei, Zijun Xia, Zhaohao Xiao, Tingshuai Yan, Liyan Yang, Yifan Yang, Zhikai Yang, Zhong Yin, Li Yuan, Liuchun Yuan, Chi Zhang, Jinyang Zhang, Junhui Zhang, Linge Zhang, Zhenyi Zhang, Zheyu Zhang, Dongjie Zhu, Hang Li, Yangang Zhang

机构 * Full author list in Contributions(完整作者列表)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Astra Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02382 2025-06-04 cs.CV cs.LG 79%

Multi-level and Multi-modal Action Anticipation

Seulgi Kim, Ghazal Kaviani, Mohit Prabhushankar, Ghassan AlRegib

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted in 2025 IEEE International Conference on Image Processing (ICIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13541 2025-06-04 cs.CV cs.LG cs.NE 79%

Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition

Jingyao Wang, Wenwen Qiang, Changwen Zheng, Fuchun Sun

机构 * University of Chinese Academy of Sciences(中国科学院大学) National Key Laboratory of Space Integrated Information System(空间信息集成系统国家重点实验室) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) Department of Computer Science and Technology(计算机科学与技术系)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17587 2025-06-03 cs.LG cs.AI 79%

Multimodal Banking Dataset: Understanding Client Needs through Event Sequences

Dzhambulat Mollaev, Alexander Kostin, Maria Postnova, Ivan Karpukhin, Ivan Kireev, Gleb Gusev, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00531 2025-06-03 cs.LG cs.AI 79%

M2WLLM: Multi-Modal Multi-Task Ultra-Short-term Wind Power Prediction Algorithm Based on Large Language Model

Hang Fana, Mingxuan Lib, Zuhan Zhanga, Long Chengc, Yujian Ye, Dunnan Liua

机构 * School of Economics and Management, North China Electric Power University(华北电力大学经济管理学院) Department of Electrical Engineering, Tsinghua University(清华大学电气工程系) School of Control and Computer Engineering, North China Electric Power University(华北电力大学控制与计算机工程学院) School of Electrical Engineering, Southeast University(东南大学电气工程学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09306 2025-06-03 cs.CV 79%

Keypoint-Integrated Instruction-Following Data Generation for Enhanced Human Pose and Action Understanding in Multimodal Models

Dewen Zhang, Wangpeng An, Hayaru Shouno

机构 * Department of Informatics, Graduate School of Informatics and Engineering, The University of Electro-Communications(信息学院、信息与工程研究生院、电通通信大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at the International Conference on Advanced Concepts for Intelligent Vision Systems (ACIVS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21374 2025-05-28 cs.CV 79%

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Junhao Cheng, Yuying Ge, Teng Wang, Yixiao Ge, Jing Liao, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG广告实验室) City University of Hong Kong(香港城市大学)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Homepage: https://github.com/TencentARC/Video-Holmes

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19125 2025-05-27 cs.CV 79%

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models

Yuqi Liu, Qin Jin, Tianyuan Qu, Xuan Liu, Yang Du, Bei Yu, Jiaya Jia

机构 * CUHK(香港中文大学) HKUST(香港理工大学) RUC(中国人民大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12746 2025-05-26 cs.AI 79%

Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs

Haruka Asanuma, Naoko Koide-Majima, Ken Nakamura, Takato Horii, Shinji Nishimoto, Masafumi Oizumi

机构 * The University of Tokyo, Graduate School of Arts and Sciences(东京大学艺术与科学研究生院) Center for Information and Neural Networks (CiNet), National Institute of Information and Communications Technology(信息与神经网络中心(CiNet),信息与通信技术国家研究所) The University of Osaka, Graduate School of Frontier Biosciences(大阪大学前沿生命科学研究生院) The University of Tokyo, Faculty of Engineering(东京大学工学部) The University of Osaka, Graduate School of Engineering Science(大阪大学工学研究院) The University of Osaka, Graduate School of Medicine(大阪大学医学研究院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12007 2025-05-23 cs.CV 79%

Multi-modal Collaborative Optimization and Expansion Network for Event-assisted Single-eye Expression Recognition

Runduo Han, Xiuping Liu, Shangxuan Yi, Yi Zhang, Hongchen Tan

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏