arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4772 篇

2503.13983 2025-04-14 cs.CV 83%

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

Jiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li, Jiannan Ge, Hongtao Xie, Yongdong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Renmin University of China(中国人民大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03735 2025-04-02 cs.CV 83%

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Chaoyu Li, Eun Woo Im, Pooyan Fazli

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23660 2025-04-01 cs.CV 83%

DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance

Junjie Zheng, Zihao Chen, Chaofan Ding, Xinhan Di

机构 * AI Lab, Giant Network(巨人网络AI实验室) University of Trento(特伦托大学)

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19406 2025-03-26 cs.CV 83%

M$^2$CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation

Ziyuan Liu, Jiawei Zhang, Wenyu Wang, Yuantao Gu

机构 * Tsinghua University(清华大学) Army Engineering University of PLA(中国人民解放军陆军工程大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19134 2025-03-26 cs.CL cs.CR 83%

MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks

Wenhao You, Bryan Hooi, Yiwei Wang, Youke Wang, Zong Ke, Ming-Hsuan Yang, Zi Huang, Yujun Cai

机构 * University of Waterloo(滑铁卢大学) National University of Singapore(新加坡国立大学) University of California, Merced(加州大学默塞德分校) University of Alberta(阿尔伯塔大学) University of Queensland(昆士兰大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13281 2025-03-25 cs.CV cs.AI cs.CL cs.MM 83%

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

Ziyang Luo, Haoning Wu, Dongxu Li, Jing Ma, Mohan Kankanhalli, Junnan Li

机构 * Salesforce AI Research(Salesforce AI研究院) Hong Kong Baptist University(香港浸会大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) The Australian National University(澳大利亚国立大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2025, Project Page: https://videoautoarena.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10523 2025-03-14 cs.CV 83%

Interactive Multimodal Fusion with Temporal Modeling

Jun Yu, Yongqi Wang, Lei Wang, Yang Zheng, Shengfan Xu

机构 * University of Science and Technology of China(中国科学技术大学) Macau University of Science and Technology(澳门科技大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17599 2025-03-14 cs.CL 83%

MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference

Zhongwei Wan, Hui Shen, Xin Wang, Che Liu, Zheda Mai, Mi Zhang

机构 * The Ohio State University(俄亥俄州立大学) Imperial College London(伦敦帝国学院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments NAACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20105 2024-12-31 cs.CV 83%

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

Jiedong Zhuang, Lu Lu, Ming Dai, Rui Hu, Jian Chen, Qiang Liu, Haoji Hu

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18060 2024-12-25 cs.CV 83%

An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLM

Wen Wen, Yilin Wang, Neil Birkbeck, Balu Adsumilli

机构 * City University of Hong Kong(香港城市大学) Google Inc.(谷歌公司)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14006 2024-12-19 cs.CV 83%

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Cong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng, Yong Liu, Zheng Zhao, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Meituan Inc.(美团公司)

专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15829 2024-12-03 cs.CV 83%

SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization

Sicheng Liu, Lintao Wang, Xiaogang Zhu, Xuequan Lu, Zhiyong Wang, Kun Hu

机构 * The University of Sydney(悉尼大学) The University of Adelaide(阿德莱德大学) La Trobe University(拉筹伯大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures, submitted to ACM Multimedia Asia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17773 2024-12-02 cs.CV 83%

XTrack: Multimodal Training Boosts RGB-X Video Object Trackers

Yuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou, Guolei Sun, Eduard Zamfi, Chao Ma, Danda Pani Paudel, Luc Van Gool, Radu Timofte

机构 * University of Wurzburg(维尔茨堡大学) INSAIT, Sofia University(保加利亚索菲亚大学INSAIT研究所) ETH Zurich(苏黎世联邦理工学院) AI Institute, Shanghai Jiao Tong University(上海交通大学AI研究院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11pages, 5figs

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09875 2024-10-15 cs.CV cs.IR 83%

ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification

Chen Mao, Chong Tan, Jingqi Hu, Min Zheng

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04955 2024-10-01 cs.CV 83%

Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations

Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang, Peng Zhai, Song Wang, Lihua Zhang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by TCSVT 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05930 2024-09-18 cs.CV 83%

MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer

Rezaul Karim, He Zhao, Richard P. Wildes, Mennatullah Siam

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

Comments Extension of CVPR'23 paper for journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10213 2024-09-17 cs.CV 83%

Neuromorphic Facial Analysis with Cross-Modal Supervision

Federico Becattini, Luca Cultrera, Lorenzo Berlincioni, Claudio Ferrari, Andrea Leonardo, Alberto Del Bimbo

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted for publication at the ECCV 2024 workshop on Neuromorphic Vision: Advantages and Applications of Event Cameras (NEVI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09362 2024-09-17 cs.CL 83%

Generating Event-oriented Attribution for Movies via Two-Stage Prefix-Enhanced Multimodal LLM

Yuanjie Lyu, Tong Xu, Zihan Niu, Bo Peng, Jing Ke, Enhong Chen

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08150 2024-09-06 cs.CV 83%

Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video Understanding

Minghui Wu, Chenxu Zhao, Anyang Su, Donglin Di, Tianyu Fu, Da An, Min He, Ya Gao, Meng Ma, Kun Yan, Ping Wang

专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ACM MULTIMEDIA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01766 2024-08-20 cs.CV 83%

MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition

Ruoyu Wang, Wenqian Wang, Jianjun Gao, Dan Lin, Kim-Hui Yap, Bingbing Li

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05523 2024-08-15 cs.HC cs.CV 83%

DeepFace-Attention: Multimodal Face Biometrics for Attention Estimation with Application to e-Learning

Roberto Daza, Luis F. Gomez, Julian Fierrez, Aythami Morales, Ruben Tolosana, Javier Ortega-Garcia

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Article accepted in the IEEE Access journal. Accessible at https://ieeexplore.ieee.org/document/10633208

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11340 2024-08-06 cs.CV cs.LG 83%

CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition

Ruoyu Wang, Chen Cai, Wenqian Wang, Jianjun Gao, Dan Lin, Wenyang Liu, Kim-Hui Yap

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09157 2024-07-15 cs.IR cs.AI cs.LG 83%

Movie Recommendation with Poster Attention via Multi-modal Transformer Feature Fusion

Linhan Xia, Yicheng Yang, Ziou Chen, Zheng Yang, Shengxin Zhu

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19654 2024-05-31 cs.AI 83%

Unlocking the Power of Spatial and Temporal Information in Medical Multimodal Pre-training

Jinxia Yang, Bing Su, Wayne Xin Zhao, Ji-Rong Wen

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.AI

Comments Accepted at ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03638 2024-05-21 cs.MM cs.HC 83%

Physical-aware Cross-modal Adversarial Network for Wearable Sensor-based Human Action Recognition

Jianyuan Ni, Hao Tang, Anne H. H. Ngu, Gaowen Liu, Yan Yan

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.MM

Comments We will be making some significant changes to the paper, including the title and methodology. We therefore wish to withdraw the paper for now

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09010 2024-04-16 cs.CV cs.LG 83%

MMA-DFER: MultiModal Adaptation of unimodal models for Dynamic Facial Expression Recognition in-the-wild

Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments accepted to CVPR 2024 ABAW Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08949 2024-04-16 cs.CL 83%

Multimodal Cross-Document Event Coreference Resolution Using Linear Semantic Transfer and Mixed-Modality Ensembles

Abhijnan Nath, Huma Jamil, Shafiuddin Rehan Ahmed, George Baker, Rahul Ghosh, James H. Martin, Nathaniel Blanchard, Nikhil Krishnaswamy

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments To appear at LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07868 2024-04-12 cs.CV 83%

MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval

Xiaojie Jin, Bowen Zhang, Weibo Gong, Kai Xu, XueQing Deng, Peng Wang, Zhao Zhang, Xiaohui Shen, Jiashi Feng

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03413 2024-04-05 cs.CV 83%

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, Mohamed Elhoseiny

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments 6 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19258 2024-03-01 cs.CV 83%

MaskFi: Unsupervised Learning of WiFi and Vision Representations for Multimodal Human Activity Recognition

Jianfei Yang, Shijie Tang, Yuecong Xu, Yunjiao Zhou, Lihua Xie

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏