arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2507.17897 2025-10-30 q-bio.NC cs.CV cs.LG 79%

Multimodal Recurrent Ensembles for Predicting Brain Responses to Naturalistic Movies (Algonauts 2025)

Semih Eren, Deniz Kucukahmetler, Nico Scherf

机构 * Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) TU Dresden(德累斯顿技术大学) School for Embedded and Composite AI (SECAI)(嵌入式与复合人工智能学院) Center for Scalable Data Analytics & AI (ScaDS.AI)(可扩展数据与人工智能中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 2 figures, 1 table. Invited report, CCN 2025 Algonauts Project session (3rd-place team). Code: https://github.com/erensemih/Algonauts2025_ModalityRNN v3: Added equal contribution footnote to author list. Corrected reference list

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25332 2025-10-30 cs.CV 79%

StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA

Yuhang Hu, Zhenyu Yang, Shihan Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Changsheng Xu

机构 * Henan Institute of Advanced Technology, Zhengzhou University(河南高级技术研究所,郑州大学) Institute of Automation, CAS(自动化研究所,中国科学院) UCAS(中国科学院大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23934 2025-10-29 cs.CY cs.AI cs.ET 79%

MFiSP: A Multimodal Fire Spread Prediction Framework

Alec Sathiyamoorthy, Wenhao Zhou, Xiangmin Zhou, Xiaodong Li, Iqbal Gondal

机构 * School of Computing Technologies, RMIT University(计算技术学院,皇家墨尔本理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03727 2025-10-29 eess.AS cs.LG 79%

Detecting Neurocognitive Disorders through Analyses of Topic Evolution and Cross-modal Consistency in Visual-Stimulated Narratives

Jinchao Li, Yuejiao Wang, Junan Li, Jiawen Kang, Bo Zheng, Ka Ho Wong, Brian Mak, Helene H. Fung, Jean Woo, Man-Wai Mak, Timothy Kwok, Vincent Mok, Xianmin Gong, Xixin Wu, Xunying Liu, Patrick C. M. Wong, Helen Meng

机构 * The Chinese University of Hong Kong(香港中文大学) The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 eess.AS

Comments 16 pages, 5 figures, accepted by "IEEE Journal of Selected Topics in Signal Processing"

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06499 2025-10-28 q-bio.NC cs.AI 79%

The ISLab Solution to the Algonauts Challenge 2025: A Multimodal Deep Learning Approach to Brain Response Prediction

Andrea Corsico, Giorgia Rigamonti, Simone Zini, Luigi Celona, Paolo Napoletano

机构 * Department of Informatics, Systems and Communication, University of Milano-Bicocca(信息学、系统与通信系,米兰-比科卡大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21406 2025-10-27 cs.CV 79%

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

Yue Feng, Jinwei Hu, Qijia Lu, Jiawei Niu, Li Tan, Shuo Yuan, Ziyi Yan, Yizhen Jia, Qingzhi He, Shiping Ge, Ethan Q. Chen, Wentong Li, Limin Wang, Jie Qin

机构 * MoE Key Laboratory of Brain-Machine Intelligence Technology, College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(脑机智能技术MoE实验室,人工智能学院,南京航空航天大学) Nanjing University(南京大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03340 2025-10-27 cs.CV 79%

Seeing the Arrow of Time in Large Multimodal Models

Zihui Xue, Mi Luo, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025, Project website: https://vision.cs.utexas.edu/projects/SeeAoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18411 2025-10-22 cs.CL cs.LG 79%

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

Yue Jiang, Jichu Li, Yang Liu, Dingkang Yang, Feng Zhou, Quyu Kong

机构 * Fudan University(复旦大学) Center for Applied Statistics and School of Statistics, Renmin University of China(应用统计中心和中国人民大学统计学院) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高级创新中心)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07447 2025-10-15 cs.CV 79%

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu

机构 * State Key Laboratory of VR Technology and Systems, School of CSE, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学计算机科学与工程学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11110 2025-10-14 cs.LG cs.AI 79%

PhysioME: A Robust Multimodal Self-Supervised Framework for Physiological Signals with Missing Modalities

Cheol-Hui Lee, Hwa-Yeon Lee, Min-Kyung Jung, Dong-Joo Kim

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10986 2025-10-14 cs.CV 79%

Mixup Helps Understanding Multimodal Video Better

Xiaoyu Ma, Ding Ding, Hao Chen

机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(计算机科学与工程学院,东南大学,南京,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其跨学科应用关键实验室(东南大学),教育部,中国)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08561 2025-10-14 cs.CV 79%

MultiCOIN: Multi-Modal COntrollable Video INbetweening

Maham Tanveer, Yang Zhou, Simon Niklaus, Ali Mahdavi Amiri, Hao Zhang, Krishna Kumar Singh, Nanxuan Zhao

机构 * Simon Fraser University(西蒙弗雷泽大学) Adobe Research(Adobe研究院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Project website: https://multicoinx.github.io/multicoin/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09266 2025-10-13 cs.CL 79%

CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation

Kaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang, Ruida Liu, Yuming Yang, Xin Xiao, Xiao Sun, Haoyang Zeng, Changzai Pan, Yidan Zhang, Jiang Zhong, Peijin Wang, Yingchao Feng

机构 * Chongqing University(重庆大学) Independent Researcher(独立研究者) University of the Chinese Academy of Sciences(中国科学院大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08143 2025-10-10 cs.CV 79%

UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution

Shian Du, Menghan Xia, Chang Liu, Quande Liu, Xintao Wang, Pengfei Wan, Xiangyang Ji

机构 * Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01674 2025-10-10 cs.CV 79%

MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs

Yipeng Du, Tiehan Fan, Kepan Nan, Rui Xie, Penghao Zhou, Xiang Li, Jian Yang, Zhenheng Yang, Ying Tai

机构 * Nanjing University(南京大学) ByteDance(字节跳动) Nankai University(南开大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15192 2025-10-08 cs.CV 79%

Leveraging Foundation Models for Multimodal Graph-Based Action Recognition

Fatemeh Ziaeetabar, Florentin Wörgötter

机构 * School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran(数学、统计与计算机科学学院,科学学院,塔里斯坦大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06461 2025-10-07 cs.CV 79%

Interactive Test-Time Adaptation with Reliable Spatial-Temporal Voxels for Multi-Modal Segmentation

Haozhi Cao, Yuecong Xu, Pengyu Yin, Xingyu Ji, Shenghai Yuan, Jianfei Yang, Lihua Xie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02287 2025-10-03 cs.CV 79%

MultiModal Action Conditioned Video Generation

Yichen Li, Antonio Torralba

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17481 2025-10-02 eess.SP cs.AI cs.LG 79%

Toward Foundational Model for Sleep Analysis Using a Multimodal Hybrid Self-Supervised Learning Framework

Cheol-Hui Lee, Hakseung Kim, Byung C. Yoon, Dong-Joo Kim

机构 * Department of Brain and Cognitive Engineering, Korea University(脑科学与认知工程系,韩国大学) Interdisciplinary Program in Precision Public Health, Korea University(精准公共卫生跨学科项目,韩国大学) Department of Radiology, Stanford University School of Medicine(放射科,斯坦福大学医学院) VA Palo Alto Health Care System(帕洛阿尔托医疗系统) Department of Neurology, Korea University College of Medicine(神经病学系,韩国大学医学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 18 pages, 5 figures

Journal ref IEEE Transactions on Cybernetics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19778 2025-10-02 cs.AI 79%

Multimodal Large Language Models for Bioimage Analysis

Shanghang Zhang, Gaole Dai, Tiejun Huang, Jianxu Chen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23085 2025-09-30 cs.IR cs.AI 79%

Enhancing Live Broadcast Engagement: A Multi-modal Approach to Short Video Recommendations Using MMGCN and User Preferences

Saeid Aghasoleymani Najafabadi, Elaheh Nabavi Nia

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19973 2025-09-26 cs.CV 79%

OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving

Pei Liu, Hongliang Lu, Haichao Liu, Haipeng Liu, Xin Liu, Ruoyu Yao, Shengbo Eben Li, Jun Ma

机构 * The Hong Kong University of Science and Technology(香港科技大学) Li Auto Inc. the School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02781 2025-09-26 q-bio.QM cs.AI cs.LG 79%

Multimodal AI predicts clinical outcomes of drug combinations from preclinical data

Yepeng Huang, Xiaorui Su, Varun Ullanat, Intae Moon, Ivy Liang, Lindsay Clegg, Damilola Olabode, Ruthie Johnson, Nicholas Ho, Megan Gibbs, Megan Gibbs, Alexander Gusev, Bino John, Marinka Zitnik

机构 * Harvard Medical School(哈佛医学院) Harvard College(哈佛学院) AstraZeneca(阿斯利康) Carnegie Mellon University(卡内基梅隆大学) Dana-Farber Cancer Institute and Harvard Medical School(达纳-法伯癌症研究所和哈佛医学院) Harvard University(哈佛大学) Broad Institute of MIT and Harvard(MIT和哈佛大学 Broad研究所) Harvard Data Science Initiative(哈佛数据科学计划)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04130 2025-09-24 cs.CV 79%

STORM: Token-Efficient Long Video Understanding for Multimodal LLMs

Jindong Jiang, Xiuyu Li, Zhijian Liu, Muyang Li, Guo Chen, Zhiqi Li, De-An Huang, Guilin Liu, Zhiding Yu, Kurt Keutzer, Sungjin Ahn, Jan Kautz, Hongxu Yin, Yao Lu, Song Han, Wonmin Byeon

机构 * NVIDIA(英伟达) Rutgers University(罗格斯大学) UC Berkeley(加州大学伯克利分校) MIT(麻省理工学院) Nanjing University(南京大学) KAIST(韩国科学技术院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17712 2025-09-23 cs.CV 79%

RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion

Geonho Bang, Minjae Seong, Jisong Kim, Geunju Baek, Daye Oh, Junhyung Kim, Junho Koh, Jun Won Choi

机构 * Seoul National University(首尔国立大学) Hanyang University(翰阳大学) Hyundai Motor Company(现代汽车公司)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15564 2025-09-23 cs.CV 79%

Show-o2: Improved Native Unified Multimodal Models

Jinheng Xie, Zhenheng Yang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室) ByteDance(字节跳动)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments NeurIPS 2025. (v3: update to include video understanding, OneIG, and more ablation study results)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15400 2025-09-22 cs.LG cs.AI cs.RO 79%

Exploring multimodal implicit behavior learning for vehicle navigation in simulated cities

Eric Aislan Antonelo, Gustavo Claudio Karl Couto, Christian Möller

机构 * Systems Engineering Department, Federal University of Santa Catarina, Florianopolis, Brazil Faculty of Science Engineering, Information Technology Åbo Akademi University, Finland

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments ENIAC conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15224 2025-09-19 cs.CV 79%

Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation

Luca Bartolomei, Enrico Mannocci, Fabio Tosi, Matteo Poggi, Stefano Mattoccia

机构 * Advanced Research Center on Electronic System (ARCES)(电子系统先进研究中心) Department of Computer Science and Engineering (DISI)(计算机科学与工程系) University of Bologna(博洛尼亚大学)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments ICCV 2025. Code: https://github.com/bartn8/depthanyevent/ Project Page: https://bartn8.github.io/depthanyevent/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13515 2025-09-18 cs.CV 79%

Multimodal Hate Detection Using Dual-Stream Graph Neural Networks

Jiangbei Yue, Shuonan Yang, Tailin Chen, Jianbo Jiao, Zeyu Fu

机构 * Multimodal Intelligence Lab, Department of Computer Science University of Exeter Exeter, UK(埃克塞特大学计算机科学系多模态智能实验室) School of Computer Science University of Birmingham Birmingham, UK(伯明翰大学计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11944 2025-09-16 cs.AI 79%

Agentic Temporal Graph of Reasoning with Multimodal Language Models: A Potential AI Aid to Healthcare

Susanta Mitra

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏