arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4757 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4757 篇

2304.11193 2026-05-14 cs.RO cs.AI cs.CV 76%

Multi-Modal World Model for Physical Robot Interactions: Simultaneous Visual and Tactile Predictions for Enhanced Accuracy

多模态世界模型用于物理机器人交互:同时进行视觉和触觉预测以提高准确性

Willow Mandil, Amir Ghalamzan-E

机构 * University of Lincoln(林肯大学) University of Sheffield(谢菲尔德大学)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.AI

AI总结 本文提出多模态世界模型,通过整合视觉和触觉信息提升机器人物理交互的预测精度,特别是在物理模糊场景中表现更优。

Comments This paper is accepted for publication in Robotics and Autonomous Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22911 2026-04-14 cs.CV cs.AI 76%

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling

ForestPrune: 通过时空森林建模实现视频多模态大语言模型的高比率视觉令牌压缩

Shaobo Ju, Baiyang Song, Tao Chen, Jiapeng Zhang, Qiong Wu, Chao Chang, HuaiXi Wang, Yiyi Zhou, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室) National University of Defense Technology(国防科技大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

AI总结 本文提出ForestPrune,一种无需训练的视频MLLM令牌压缩方法,通过时空森林建模实现高效高比率压缩,实验表明其在视频任务中效果优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16788 2026-03-16 cs.CV cs.AI 76%

Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation

迈向可解释AI:基于视频的图像描述生成的多模态Transformer

Lakshita Agarwal, Bindu Verma

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.AI

AI总结 本文提出一种多模态Transformer框架,结合文本和视觉模态生成视频描述,通过ResNet50提取视频特征并结合GPT-2模型,提升描述质量和可解释性。

Journal ref 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01817 2025-07-23 cs.CV cs.AI cs.CY 76%

Aligning AI with Public Values: Deliberation and Decision-Making for Governing Multimodal LLMs in Political Video Analysis

Tanusree Sharma, Yujin Potter, Zachary Kilhoffer, Yun Huang, Dawn Song, Yang Wang

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21116 2025-07-09 cs.CV cs.AI 76%

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

Yujia Liang, Jile Jiao, Xuetao Feng, Zixuan Ye, Yuan Wang, Zhicheng Wang

机构 * School of AIA, Huazhong University of Science and Technology(华中科技大学人工智能学院) Deepeleph Intelligent Technology(DeepEleph智能技术) JD Explore Academy(JD探索学院) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07804 2025-03-06 eess.IV cs.AI cs.CV 76%

XLSTM-HVED: Cross-Modal Brain Tumor Segmentation and MRI Reconstruction Method Using Vision XLSTM and Heteromodal Variational Encoder-Decoder

Shenghao Zhu, Yifei Chen, Shuo Jiang, Weihong Chen, Chang Liu, Yuanhan Wang, Xu Chen, Yifan Ke, Feiwei Qin, Changmiao Wang, Zhu Zhu

专题命中 视频多模态 :cross-modal(title);分类 cs.CV、cs.AI

Comments 5 pages, 2 figures

Journal ref ISBI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16471 2025-01-29 cs.LG cs.AI eess.AS eess.IV q-bio.NC 76%

SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching Experiments

Simon Dahan, Gabriel Bénédict, Logan Z. J. Williams, Yourong Guo, Daniel Rueckert, Robert Leech, Emma C. Robinson

专题命中 视频多模态 :multimodal(title);分类 cs.AI、eess.AS

Comments 27 pages, accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01422 2025-01-03 cs.CV cs.AI cs.LG 76%

Multi-Modal Video Feature Extraction for Popularity Prediction

Haixu Liu, Wenning Wang, Haoxiang Zheng, Penghao Jiang, Qirui Wang, Ruiqing Yan, Qiuzhuang Sun

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.AI

Comments INFORMS 2024 Data Challenge Competition

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06712 2024-12-10 cs.LG cs.CL cs.CV 76%

How to Merge Your Multimodal Models Over Time?

Sebastian Dziadzio, Vishaal Udandarao, Karsten Roth, Ameya Prabhu, Zeynep Akata, Samuel Albanie, Matthias Bethge

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL

Comments Technical Report. Code at https://github.com/ExplainableML/fomo_in_flux

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03531 2024-11-07 cs.CV cs.AI 76%

Personalized Video Summarization by Multimodal Video Understanding

Brian Chen, Xiangyuan Zhao, Yingnan Zhu

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments In Proceedings of CIKM 2024 Applied Research Track

Journal ref 33rd ACM International Conference on Information and Knowledge Management (CIKM 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.10282 2024-10-28 cs.CV cs.AI cs.LG eess.IV 76%

ChiNet: Deep Recurrent Convolutional Learning for Multimodal Spacecraft Pose Estimation

Duarte Rondao, Nabil Aouf, Mark A. Richardson

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03340 2024-08-08 cs.MM cs.CV 76%

An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval

Mahesh Kandhare, Thibault Gisselbrecht

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.MM

Comments 19 pages, 24 figures (65 images)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09818 2024-05-27 cs.CL cs.AI 76%

SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models

Lee Hyun, Kim Sung-Bin, Seungju Han, Youngjae Yu, Tae-Hyun Oh

专题命中 视频多模态 :multimodal(title);分类 cs.CL、cs.AI

Comments 19 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10586 2024-04-30 cs.CV cs.CL 76%

VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools

Ji Qi, Kaixuan Ji, Jifan Yu, Duokang Wang, Bin Xu, Lei Hou, Juanzi Li

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10790 2024-04-18 cs.CR cs.AI cs.CV cs.LG 76%

Multimodal Attack Detection for Action Recognition Models

Furkan Mumcu, Yasin Yilmaz

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04479 2023-12-08 cs.CV cs.AI 76%

GSGFormer: Generative Social Graph Transformer for Multimodal Pedestrian Trajectory Prediction

Zhongchang Luo, Marion Robin, Pavan Vasishta

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13372 2023-10-17 cs.CV cs.AI cs.LG 76%

Localizing Moments in Long Video Via Multimodal Guidance

Wayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo, Fabian Caba Heilbron, Bernard Ghanem

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10505 2023-04-21 cs.CV cs.AI cs.LG 76%

Video Pre-trained Transformer: A Multimodal Mixture of Pre-trained Experts

Kastan Day, Daniel Christl, Rohan Salvi, Pranav Sriram

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05991 2023-02-22 cs.CV cs.CL cs.LG 76%

Text-Derived Knowledge Helps Vision: A Simple Cross-modal Distillation for Video-based Action Anticipation

Sayontan Ghosh, Tanvi Aggarwal, Minh Hoai, Niranjan Balasubramanian

专题命中 视频多模态 :cross-modal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.03977 2022-05-26 eess.AS cs.LG cs.MM stat.ML 76%

Multimodal active speaker detection and virtual cinematography for video conferencing

Ross Cutler, Ramin Mehran, Sam Johnson, Cha Zhang, Adam Kirk, Oliver Whyte, Adarsh Kowdle

专题命中 视频多模态 :multimodal(title);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.17066 2022-05-20 eess.SP cs.AI cs.CV cs.LG 76%

Cross-modal Learning of Graph Representations using Radar Point Cloud for Long-Range Gesture Recognition

Souvik Hazra, Hao Feng, Gamze Naz Kiprit, Michael Stephan, Lorenzo Servadei, Robert Wille, Robert Weigel, Avik Santra

专题命中 视频多模态 :cross-modal(title);分类 cs.CV、cs.AI

Comments Accepted by IEEE Sensor Array and Multichannel Signal Processing Workshop (SAM 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.09276 2021-11-18 cs.CV cs.CL 76%

Induce, Edit, Retrieve: Language Grounded Multimodal Schema for Instructional Video Retrieval

Yue Yang, Joongwon Kim, Artemis Panagopoulou, Mark Yatskar, Chris Callison-Burch

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.01173 2021-02-03 cs.LG cs.AI cs.MM 76%

Multi-modal Ensemble Models for Predicting Video Memorability

Tony Zhao, Irving Fang, Jeffrey Kim, Gerald Friedland

专题命中 视频多模态 :multi-modal(title);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02469 2021-01-08 cs.CV cs.AI 76%

Multimodal Gait Recognition for Neurodegenerative Diseases

Aite Zhao, Jianbo Li, Junyu Dong, Lin Qi, Qianni Zhang, Ning Li, Xin Wang, Huiyu Zhou

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04794 2020-12-10 cs.CV cs.AI cs.RO 76%

Deep Learning based Multi-Modal Sensing for Tracking and State Extraction of Small Quadcopters

Zhibo Zhang, Chen Zeng, Maulikkumar Dhameliya, Souma Chowdhury, Rahul Rai

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09747 2020-08-25 cs.CV cs.AI cs.LG 76%

Towards Improved Human Action Recognition Using Convolutional Neural Networks and Multimodal Fusion of Depth and Inertial Sensor Data

Zeeshan Ahmad, Naimul Khan

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.05187 2019-11-14 cs.CV cs.LG cs.SD eess.AS stat.ML 76%

AI in Pursuit of Happiness, Finding Only Sadness: Multi-Modal Facial Emotion Recognition Challenge

Carl Norman

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.08859 2019-09-20 cs.CL cs.CV 76%

Procedural Reasoning Networks for Understanding Multimodal Procedures

Mustafa Sercan Amac, Semih Yagcioglu, Aykut Erdem, Erkut Erdem

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL

Comments Accepted to CoNLL 2019. The project website with code and demo is available at https://hucvl.github.io/prn/

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.03466 2019-09-10 cs.CV cs.AI cs.HC cs.LG 76%

Multi-Modal Three-Stream Network for Action Recognition

Muhammad Usman Khalid, Jie Yu

专题命中 视频多模态 :multi-modal(title);分类 cs.CV、cs.AI

Comments Presented in IEEE ICPR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.04545 2018-08-15 cs.LG cs.AI cs.CV cs.GR stat.ML 76%

MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics

Xinchen Yan, Akash Rastogi, Ruben Villegas, Kalyan Sunkavalli, Eli Shechtman, Sunil Hadap, Ersin Yumer, Honglak Lee

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments Published at ECCV 2018

详情

展开后加载摘要…

URL PDF HTML 收藏