arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2408.11286 2024-08-23 cs.CV 79%

Video Emotion Open-vocabulary Recognition Based on Multimodal Large Language Model

Mengying Ge, Dongkai Tang, Mingyang Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01165 2024-08-13 cs.CL 79%

LITE: Modeling Environmental Ecosystems with Multimodal Large Language Models

Haoran Li, Junqi Liu, Zexian Wang, Shiyuan Luo, Xiaowei Jia, Huaxiu Yao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments COLM camera-ready version. Code is released at https://github.com/hrlics/LITE

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16123 2024-07-24 cs.LG cs.AI 79%

Towards Effective Fusion and Forecasting of Multimodal Spatio-temporal Data for Smart Mobility

Chenxing Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13137 2024-07-19 cs.CV 79%

OE-BevSeg: An Object Informed and Environment Aware Multimodal Framework for Bird's-eye-view Vehicle Semantic Segmentation

Jian Sun, Yuqi Dai, Chi-Man Vong, Qing Xu, Shengbo Eben Li, Jianqiang Wang, Lei He, Keqiang Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12798 2024-07-19 cs.CV 79%

Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval

Wenjun Li, Shudong Wang, Dong Zhao, Shenghui Xu, Zhaoming Pan, Zhimin Zhang

专题命中 视频多模态 :multi-modal(title);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19917 2024-07-17 cs.CV 79%

Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition

Masashi Hatano, Ryo Hachiuma, Ryo Fujii, Hideo Saito

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ECCV'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11481 2024-07-16 cs.CV 79%

VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Yue Fan, Xiaojian Ma, Rujie Wu, Yuntao Du, Jiaqi Li, Zhi Gao, Qing Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ECCV-24; Project page: videoagent.github.io; First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07775 2024-07-15 cs.RO cs.AI 79%

Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Hao-Tien Lewis Chiang, Zhuo Xu, Zipeng Fu, Mithun George Jacob, Tingnan Zhang, Tsang-Wei Edward Lee, Wenhao Yu, Connor Schenck, David Rendleman, Dhruv Shah, Fei Xia, Jasmine Hsu, Jonathan Hoech, Pete Florence, Sean Kirmani, Sumeet Singh, Vikas Sindhwani, Carolina Parada, Chelsea Finn, Peng Xu, Sergey Levine, Jie Tan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05419 2024-07-09 cs.CV cs.IR 79%

Multimodal Language Models for Domain-Specific Procedural Video Summarization

Nafisa Hussain

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02218 2024-07-08 cs.CV 79%

Multi-Modal Video Dialog State Tracking in the Wild

Adnen Abdessaied, Lei Shi, Andreas Bulling

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14098 2024-07-08 cs.CV 79%

HeartBeat: Towards Controllable Echocardiography Video Synthesis with Multimodal Conditions-Guided Diffusion Models

Xinrui Zhou, Yuhao Huang, Wufeng Xue, Haoran Dou, Jun Cheng, Han Zhou, Dong Ni

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03836 2024-07-08 cs.CV cs.LG 79%

ADAPT: Multimodal Learning for Detecting Physiological Changes under Missing Modalities

Julie Mordacq, Leo Milecki, Maria Vakalopoulou, Steve Oudot, Vicky Kalogeiton

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at MIDL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12235 2024-07-02 cs.CV 79%

Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM

Huaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo, Chuchu Han, Xiaonan Huang, Changxin Gao, Yuehuan Wang, Nong Sang

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 19 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15556 2024-06-25 cs.CV 79%

Open-Vocabulary Temporal Action Localization using Multimodal Guidance

Akshita Gupta, Aditya Arora, Sanath Narayan, Salman Khan, Fahad Shahbaz Khan, Graham W. Taylor

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09781 2024-06-17 cs.CV 79%

GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

Yiqi Wu, Xiaodan Hu, Ziming Fu, Siling Zhou, Jiangong Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09711 2024-06-17 cs.CV 79%

AnimalFormer: Multimodal Vision Framework for Behavior-based Precision Livestock Farming

Ahmed Qazi, Taha Razzaq, Asim Iqbal

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09076 2024-06-14 cs.CL 79%

3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection

Thye Shan Ng, Feiqi Cao, Soyeon Caren Han

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05496 2024-06-11 cs.CL 79%

Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities

Sai Munikoti, Ian Stewart, Sameera Horawalavithana, Henry Kvinge, Tegan Emerson, Sandra E Thompson, Karl Pazdernik

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments 25 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04575 2024-06-10 cs.LG cs.AI stat.AP stat.ML 79%

Optimization of geological carbon storage operations with multimodal latent dynamic model and deep reinforcement learning

Zhongzheng Wang, Yuntian Chen, Guodong Chen, Dongxiao Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04236 2024-06-07 cs.CV 79%

Understanding Information Storage and Transfer in Multi-modal Large Language Models

Samyadeep Basu, Martin Grayson, Cecily Morrison, Besmira Nushi, Soheil Feizi, Daniela Massiceti

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15813 2024-05-28 cs.CV 79%

From CNNs to Transformers in Multimodal Human Action Recognition: A Survey

Muhammad Bilal Shaikh, Syed Mohammed Shamsul Islam, Douglas Chai, Naveed Akhtar

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 5 figures and 3 Tables. To appear in ACM Trans. Multimedia Comput. Commun. Appl.(TOMM) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12708 2024-05-22 cs.CV 79%

Multimodal video analysis for crowd anomaly detection using open access tourism cameras

Alejandro Dionis-Ros, Joan Vila-Francés, Rafael Magdalena-Benedicto, Fernando Mateo, Antonio J. Serrano-López

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11776 2024-05-20 cs.LG cs.CV 79%

3D object quality prediction for Metal Jet Printer with Multimodal thermal encoder

Rachel, Chen, Wenjia Zheng, Sandeep Jalui, Pavan Suri, Jun Zeng

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03272 2024-05-07 cs.CV 79%

WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Yuanhan Zhang, Kaichen Zhang, Bo Li, Fanyi Pu, Christopher Arif Setiadharma, Jingkang Yang, Ziwei Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03091 2024-05-07 cs.CV cs.LG 79%

Research on Image Recognition Technology Based on Multimodal Deep Learning

Jinyin Wang, Xingchen Li, Yixuan Jin, Yihao Zhong, Keke Zhang, Chang Zhou

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14471 2024-04-29 cs.CV 79%

Narrative Action Evaluation with Prompt-Guided Multimodal Interaction

Shiyi Zhang, Sule Bai, Guangyi Chen, Lei Chen, Jiwen Lu, Junle Wang, Yansong Tang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05726 2024-04-25 cs.CV 79%

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, Ser-Nam Lim

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2024. Project Page https://boheumd.github.io/MA-LMM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11470 2024-04-18 cs.CV 79%

Exploring Missing Modality in Multimodal Egocentric Datasets

Merey Ramazanova, Alejandro Pardo, Humam Alwassel, Bernard Ghanem

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09276 2024-04-18 cs.CV 79%

Transformer-based Multimodal Change Detection with Multitask Consistency Constraints

Biyuan Liu, Huaixin Chen, Kun Li, Michael Ying Yang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10429 2024-04-17 cs.AI 79%

MEEL: Multi-Modal Event Evolution Learning

Zhengwei Tao, Zhi Jin, Junqiang Huang, Xiancai Chen, Xiaoying Bai, Haiyan Zhao, Yifan Zhang, Chongyang Tao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏