arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2511.07290 2025-11-11 eess.IV cs.CV cs.MM 81%

CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video

Xinyi Wang, Angeliki Katsenou, Junxiao Shen, David Bull

机构 * School of Computer Science, University of Bristol(布里斯托大学计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14809 2025-11-05 cs.CV cs.MM cs.RO 81%

Light Future: Multimodal Action Frame Prediction via InstructPix2Pix

Zesen Zhong, Duomin Zhang, Yijia Li

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳))

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 9 pages including appendix, 4 tables, 8 figures, to be submitted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17394 2025-10-28 cs.CV cs.AI 81%

HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs

Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin

机构 * Xinjiang University(新疆大学) Xi'an Jiaotong University(西安交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21761 2025-10-28 cs.RO cs.AI cs.CV 81%

J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino

机构 * Division of Information Science, NAIST(NAIST信息科学系) Guardian Robot Project, RIKEN(RIKEN守护机器人项目)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17038 2025-10-21 cs.RO cs.AI cs.CV 81%

DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation

Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi

机构 * Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学) The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院)) Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08559 2025-10-10 cs.CV cs.AI 81%

SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models

Andong Deng, Taojiannan Yang, Shoubin Yu, Lincoln Spencer, Mohit Bansal, Chen Chen, Serena Yeung-Levy, Xiaohan Wang

机构 * University of Central Florida(中央佛罗里达大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02790 2025-10-06 cs.CV cs.CL 81%

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

Xiangfeng Wang, Xiao Li, Yadong Wei, Xueyu Song, Yang Song, Xiaoqiang Xia, Fangrui Zeng, Zaiyi Chen, Liu Liu, Gu Xu, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学) ByteDance China(字节跳动中国)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03569 2025-10-06 cs.CV cs.AI 81%

Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement

Bing Li, Jiaxin Chen, Dongming Zhang, Xiuguo Bao, Di Huang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to IJCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25393 2025-10-02 cs.CV cs.AI 81%

Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction

Wendong Yao, Binhua Huang, Soumyabrata Dev

机构 * ADAPT SFI Research Centre, School of Computer Science, University College Dublin(ADAPT SFI研究所以及计算机科学学院,都柏林大学学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments This paper is submitted to IEEE Transactions on Geoscience and Remote Sensing for reviewing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07032 2025-10-01 cs.CL cs.CV 81%

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin'ichi Satoh, Michael Felsberg, Mubarak Shah, Salman Khan, Fahad Shahbaz Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫德赫·本·扎耶德人工智能大学) University of Central Florida(中央佛罗里达大学) Islamic University of Technology(伊斯兰技术大学) Air University(空军大学) ETH Zurich(苏黎世联邦理工学院) Technische Universität München(慕尼黑技术大学) National Institute of Informatics(国家信息研究所) Australian National University(澳大利亚国立大学) Linköping University(利尔贝里大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23044 2025-09-30 cs.CV cs.AI 81%

MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition

Ye-eun Kim, Suhyeon Lim, Andrew J. Choi

机构 * National Rehabilitation Center, Ministry of Health and Welfare, Korea(韩国卫生福利部国家康复中心) Gachon University(高丽大学)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18436 2025-09-30 cs.AI cs.CL cs.DB 81%

Memory-QA: Answering Recall Questions Based on Multimodal Memories

Hongda Jiang, Xinyuan Zhang, Siddhant Garg, Rishab Arora, Shiun-Zu Kuo, Jiayang Xu, Ankur Bansal, Christopher Brossman, Yue Liu, Aaron Colak, Ahmed Aly, Anuj Kumar, Xin Luna Dong

机构 * Meta Reality Labs(Meta现实实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22014 2025-09-29 cs.CV cs.AI cs.HC cs.RO 81%

Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics

Saurav Jha, Stefan K. Ehrlich

机构 * SETLabs Resarch GmbH(SETL实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13763 2025-09-16 cs.CV cs.AI 81%

Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models

Zhawnen Chen, Tianchun Wang, Yizhou Wang, Michal Kosinski, Xiang Zhang, Yun Fu, Sheng Li

机构 * University of Virginia(弗吉尼亚大学) The Pennsylvania State University(宾夕法尼亚州立大学) Northeastern University(东北大学) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01177 2025-09-03 cs.CV cs.AI cs.HC eess.SP 81%

DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion

Junxiang Liu, Junming Lin, Jiangtong Li, Jie Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00357 2025-09-03 cs.CV cs.AI cs.LG 81%

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu, Danny T. M. Chan, Nassir Navab, Hongbin Liu, Zhen Lei, Jiebo Luo

机构 * Hong Kong Institute of Science & Innovation(香港科学与工业创新研究院) CAMP, Technische Universität München(CAMP,慕尼黑技术大学) Department of Surgery, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院外科部)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09362 2025-08-14 cs.CV cs.AI cs.LG 81%

FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition

Md. Milon Islam, Md Rezwanul Haque, S M Taslim Uddin Raju, Fakhri Karray

机构 * University of Waterloo(滑铁卢大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted for the IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, Hawaii, USA. 1st MSLR Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04900 2025-08-08 cs.CV cs.AI 81%

Revealing Temporal Label Noise in Multimodal Hateful Video Classification

Shuonan Yang, Tailin Chen, Rahul Singh, Jiangbei Yue, Jianbo Jiao, Zeyu Fu

机构 * Multimodal Intelligence Lab(多模态智能实验室) Department of Computer Science(计算机科学系) University of Exeter(埃克塞特大学) University of Birmingham(伯明翰大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05185 2025-08-05 q-fin.CP cs.AI cs.MM 81%

Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance

Fengbin Zhu, Junfeng Li, Liangming Pan, Wenjie Wang, Fuli Feng, Chao Wang, Huanbo Luan, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Peking University(北京大学) University of Science and Technology of China(中国科学技术大学) Estates Pte Ltd(6Estates私人有限公司)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments Accepted by MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21161 2025-07-30 cs.CV cs.AI cs.LG 81%

Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues

Pallavi Zambare, Venkata Nikhil Thanikella, Ying Liu

机构 * Departmrnt of computer science(计算机科学系) Texas Tech University(得克萨斯科技大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in IEEE 3rd International Conference on Artificial Intelligence, Blockchain, and Internet of Things (AIBThings 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18252 2025-07-25 cs.HC cs.AI cs.CL cs.LG 81%

Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning

Dongyang Guo, Yasmeen Abdrabou, Enkeleda Thaqi, Enkelejda Kasneci

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14766 2025-07-22 cs.LG cs.AI cs.CV 81%

CXR-TFT: Multi-Modal Temporal Fusion Transformer for Predicting Chest X-ray Trajectories

Mehak Arora, Ayman Ali, Kaiyuan Wu, Carolyn Davis, Takashi Shimazui, Mahmoud Alwakeel, Victor Moas, Philip Yang, Annette Esper, Rishikesan Kamaleswaran

机构 * Duke University(杜克大学) Emory University(埃默里大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments In Review for MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05939 2025-07-09 cs.CL cs.MM 81%

Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors

Bing Wang, Ximing Li, Mengzhe Ye, Changchun Li, Bo Fu, Jianfeng Qu, Lin Yuanbo Wu

机构 * College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院) College of Software, Jilin University(吉林大学软件学院) School of Computer and Artificial Intelligence, Liaoning Normal University(辽宁师范大学计算机与人工智能学院) School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Department of Computer Science, Swansea University(斯旺西大学计算机科学系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted by ACM MM 2025. 10 pages, 6 figures. Code: https://github.com/wangbing1416/DAEDCMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02904 2025-07-08 cs.CV cs.AI 81%

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis

Charlton Teo

机构 * Department of Computer Science(计算机科学系) School of Computing(计算学院) National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments B.Comp. dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15220 2025-06-18 cs.CV cs.AI 81%

Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Kun Yuan, Vinkle Srivastav, Tong Yu, Joel L. Lavanchy, Jacques Marescaux, Pietro Mascagni, Nassir Navab, Nicolas Padoy

机构 * University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学,法国国家科学研究中心,法国国家卫生研究院,ICube,UMR7357,法国斯特拉斯堡) University Digestive Health Care Center – Clarunis, 4002 Basel, Switzerland(消化健康研究中心–Clarunis,瑞士巴塞尔)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by Medical Image Analysis (MedIA), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13322 2025-06-17 cs.CV cs.AI 81%

Active Multimodal Distillation for Few-shot Action Recognition

Weijia Feng, Yichen Zhu, Ruojia Zhang, Chenyang Wang, Fei Ma, Xiaobao Wang, Xiaobai Li

机构 * College of Computer and Information Engineering, Tianjin Normal University(天津师范大学计算机与信息工程学院) College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳)) College of Intelligence and Computing, Tianjin University(天津大学智能科学与计算学院) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江省区块链与数据安全国家重点实验室) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Hangzhou(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments IJCAI 2025, the 34th International Joint Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10415 2025-06-13 cs.CL cs.CV 81%

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?

Yingjin Song, Yupei Du, Denis Paperno, Albert Gatt

机构 * Utrecht University(乌特雷赫大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 27 pages, 14 figures. Accepted to ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09081 2025-06-11 cs.CV cs.AI 81%

Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment

Xiaowei Bi, Zheyuan Xu

机构 * Northwestern University(西北大学) IEEE Member(IEEE会员)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01757 2025-06-03 cs.CV cs.AI 81%

Efficient Egocentric Action Recognition with Multimodal Data

Marco Calzavara, Ard Kastrati, Matteo Macchini, Dushan Vasilevski, Roger Wattenhofer

机构 * ETH Zurich(苏黎世联邦理工学院) Magic Leap(Magic Leap公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted as an extended abstract at the Second Joint Egocentric Vision (EgoVis) Workshop, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02406 2025-05-29 cs.CV cs.AI cs.DC cs.LG 81%

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models

Tzu-Tao Chang, Shivaram Venkataraman

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏