arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2505.16630 2025-05-23 cs.CV cs.AI 81%

SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding

Sushant Gautam, Cise Midoglu, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen, Mubarak Shah

机构 * SimulaMet and OsloMet, Norway(SimulaMet和OsloMet,挪威) Forzasys, Norway(Forzasys,挪威) SimulaMet, Norway(SimulaMet,挪威) SimulaMet, OsloMet, and Forzasys, Norway(SimulaMet、OsloMet和Forzasys,挪威) University of Central Florida, USA(佛罗里达中央大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16279 2025-05-23 cs.MM cs.CV 81%

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

Junjie Zheng, Zihao Chen, Chaofan Ding, Yunming Liang, Yihan Fan, Huan Yang, Lei Xie, Xinhan Di

机构 * Giant NetworkChina(巨网中国) The East China University of Science and TechnologyChina(东华大学) Northwestern Polytechnical UniversityChina(西北工业大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments 5 pages, 4 figures, accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12368 2025-05-21 cs.CV cs.CL 81%

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenxi Wu, Yubo Ma, Haodong Duan, Wenwei Zhang, Kai Chen, Dahua Lin, Jiaqi Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Nanjing University(南京大学) Fudan University(复旦大学) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11083 2025-05-21 cs.CV cs.CL 81%

Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning

Xiaohao Xu, Yunkang Cao, Huaxin Zhang, Nong Sang, Xiaonan Huang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Best Student Paper Award at IEEE International Conference on Computer Supported Cooperative Work in Design, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11926 2025-05-20 cs.CV cs.AI 81%

SafeVid: Toward Safety Aligned Video Large Multimodal Models

Yixu Wang, Jiaxin Song, Yifeng Gao, Xin Wang, Yang Yao, Yan Teng, Xingjun Ma, Yingchun Wang, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10836 2025-05-19 cs.CL cs.CV 81%

Multimodal Event Detection: Current Approaches and Defining the New Playground through LLMs and VLMs

Abhishek Dey, Aabha Bothera, Samhita Sarikonda, Rishav Aryan, Sanjay Kumar Podishetty, Akshay Havalgi, Gaurav Singh, Saurabh Srivastava

机构 * George Mason University(乔治·马歇尔大学) Amazon(亚马逊公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at NLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01680 2025-05-06 cs.CV cs.AI cs.HC math.PR 81%

Automated ARAT Scoring Using Multimodal Video Analysis, Multi-View Fusion, and Hierarchical Bayesian Models: A Clinician Study

Tamim Ahmed, Thanassis Rikakis

机构 * University of Southern California(南加州大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19918 2025-04-29 cs.CV cs.AI 81%

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI

Hugo Georgenthum, Cristian Cosentino, Fabrizio Marozzo, Pietro Liò

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17213 2025-04-29 cs.CV cs.AI 81%

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding

Shiwen Cao, Zhaoxing Zhang, Junming Jiao, Juyi Qiao, Guowen Song, Rong Shen, Xiangbing Meng

机构 * Li Auto Inc.(利亚汽车公司) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14432 2025-04-22 cs.CV cs.AI 81%

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task

Ahmad Khalil, Mahmoud Khalil, Alioune Ngom

机构 * University of Windsor(温莎大学)

专题命中 视频多模态 :multi-modal(title);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14429 2025-04-22 cs.CV cs.AI 81%

ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations

Ahmad Khalil, Mahmoud Khalil, Alioune Ngom

机构 * University of Windsor(温莎大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11493 2025-04-17 cs.RO cs.AI cs.CV 81%

Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning

Azizul Zahid, Jie Fan, Farong Wang, Ashton Dy, Sai Swaminathan, Fei Liu

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments ICRA'25 Workshop: Human-Centered Robot Learning in the Era of Big Data and Large Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11232 2025-04-16 cs.CV cs.MM 81%

Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset

Elisa Ancarani, Julie Tores, Lucile Sassatelli, Rémy Sun, Hui-Yin Wu, Frédéric Precioso

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 6 pages, 8 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05030 2025-04-08 cs.CV cs.MM 81%

AsyReC: A Multimodal Graph-based Framework for Spatio-Temporal Asymmetric Dyadic Relationship Classification

Wang Tang, Fethiye Irmak Dogan, Linbo Qing, Hatice Gunes

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04550 2025-04-08 cs.CV cs.AI cs.LG cs.RO 81%

Advancing Egocentric Video Question Answering with Multimodal Large Language Models

Alkesh Patel, Vibhav Chitalia, Yinfei Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22233 2025-04-01 cs.CV cs.AI cs.IR 81%

ContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising

Ashutosh Chaubey, Anoubhav Agarwaal, Sartaki Sinha Roy, Aayush Agrawal, Susmita Ghose

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published at WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22610 2025-03-31 cs.HC cs.AI cs.CL cs.CY cs.LG 81%

Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users

Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05421 2025-03-21 cs.CV cs.AI 81%

EPAM-Net: An Efficient Pose-driven Attention-guided Multimodal Network for Video Action Recognition

Ahmed Abdelkawy, Asem Ali, Aly Farag

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref Neurocomputing, Volume 633, 7 June 2025, 129781

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07667 2025-03-21 cs.LG cs.AI cs.CV eess.SP 81%

CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

Wei Dai, Peilin Chen, Malinda Lu, Daniel Li, Haowen Wei, Hejie Cui, Paul Pu Liang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13531 2025-03-19 cs.CV cs.AI cs.CY cs.LG 81%

Context-aware Multimodal AI Reveals Hidden Pathways in Five Centuries of Art Evolution

Jin Kim, Byunghwee Lee, Taekho You, Jinhyuk Yun

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 30 pages, 4 figures. Some example paintings are blurred to avoid potential copyright violations

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12077 2025-03-18 cs.CV cs.AI 81%

V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents

Zhengrong Yue, Shaobin Zhuang, Kunchang Li, Yanbo Ding, Yali Wang

专题命中 视频多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00064 2025-03-10 cs.LG cs.AI cs.CV cs.RO 81%

M2Distill: Multi-Modal Distillation for Lifelong Imitation Learning

Kaushik Roy, Akila Dissanayake, Brendan Tidd, Peyman Moghadam

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments IEEE ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13130 2025-02-19 cs.CV cs.AI cs.HC cs.LG cs.RO 81%

Magma: A Foundation Model for Multimodal AI Agents

Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, Yuquan Deng, Lars Liden, Jianfeng Gao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 29 pages, 16 figures, technical report from MSR

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10674 2025-02-19 cs.CV cs.CL 81%

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

Mohamed Fazli Imam, Chenyang Lyu, Alham Fikri Aji

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Our dataset can be found at \url{https://huggingface.co/datasets/fazliimam/temporal-vqa}

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19100 2025-02-18 cs.CV cs.AI 81%

VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

Lawrence Jang, Yinheng Li, Dan Zhao, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, Kazuhito Koishida

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10145 2025-02-17 cs.CV cs.MM 81%

Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling

Xinyu Li, Marwa Mahmoud

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03857 2025-02-07 cs.LG cs.CL cs.CV 81%

MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition

Stefan Gerd Fritsch, Cennet Oguz, Vitor Fortes Rey, Lala Ray, Maximilian Kiefer-Emmanouilidis, Paul Lukowicz

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16786 2025-01-29 cs.CV cs.CL 81%

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding

Yun Li, Zhe Liu, Yajing Kong, Guangrui Li, Jiyuan Zhang, Chao Bian, Feng Liu, Lina Yao, Zhenbang Sun

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15438 2025-01-28 cs.CV cs.MM 81%

Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection

Han Wang, Rui Yang Tan, Roy Ka-Wei Lee

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments 10 pages, 4 figures, THE WEB CONFERENCE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09355 2025-01-17 cs.AI cs.CV cs.ET cs.MA 81%

YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks

Saptarashmi Bandyopadhyay, Vikas Bahirwani, Lavisha Aggarwal, Bhanu Guda, Lin Li, Andrea Colaco

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏