arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4729 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4729 篇

2506.16450 2025-06-23 cs.CV 83%

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?

Giuseppe Lando, Rosario Forte, Giovanni Maria Farinella, Antonino Furnari

机构 * Department of Mathematics and Computer Science(数学与计算机科学系)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10430 2025-06-13 cs.CV 83%

MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment

Shuo wang, Jihao Zhang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09668 2025-06-12 cs.CV cs.LG 83%

CINeMA: Conditional Implicit Neural Multi-Modal Atlas for a Spatio-Temporal Representation of the Perinatal Brain

Maik Dannecker, Vasiliki Sideri-Lampretsa, Sophie Starck, Angeline Mihailov, Mathieu Milh, Nadine Girard, Guillaume Auzias, Daniel Rueckert

机构 * School of Computation, Information and Technology, and the School of Medicine and Health(计算信息学院及医学健康学院) Technical University of Munich(慕尼黑技术大学) Department of Computing(计算系) Institut de Neurosciences de la Timone, UMR 7289, CNRS, Aix-Marseille Université(神经科学研究所,UMR 7289,CNRS,艾克斯-马赛大学) Aix-Marseille Univ, APHM, Service de Neuroradiologie Diagnostique et Interventionnelle, Hôpital de la Timone(艾克斯-马赛大学,APHM,诊断与介入神经放射科,泰米翁医院) Aix-Marseille Univ, APHM, service de neurologie pédiatrique, Hôpital de la Timone(艾克斯-马赛大学,APHM,儿童神经科,泰米翁医院)

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Work currently under revision for IEEE TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08649 2025-06-11 cs.CV 83%

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization

Zhiyi Zhu, Xiaoyu Wu, Youwei Lu

机构 * Department of Information and Communication Engineering(信息与通信工程系)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08466 2025-06-10 cs.CV 83%

Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

Quan Zhang, Jinwei Fang, Rui Yuan, Xi Tang, Yuxin Qi, Ke Zhang, Chun Yuan

机构 * Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to CVPR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00304 2025-06-04 cs.CV 83%

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

Yuxiang Guo, Faizan Siddiqui, Yang Zhao, Rama Chellappa, Shao-Yuan Lo

机构 * Johns Hopkins University(约翰霍普金斯大学) Honda Research Institute USA(本田研究院美国)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Paper is accepted by IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08282 2025-06-03 cs.CV 83%

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Hongyu Li, Jinyu Chen, Ziyu Wei, Shaofei Huang, Tianrui Hui, Jialin Gao, Xiaoming Wei, Si Liu

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Meituan(美团)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24476 2025-06-02 cs.CV 83%

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model

Yuting Zhang, Hao Lu, Qingyong Hu, Yin Wang, Kaishen Yuan, Xin Liu, Kaishun Wu

机构 * The Hong Kong University of Science & Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science & Technology(香港科技大学) Zhejiang University(浙江大学) Lappeenranta-Lahti University of Technology(拉佩兰塔-拉赫蒂技术大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19400 2025-05-27 cs.AI cs.CL cs.CV cs.MM 83%

TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding

Max Ku, Thomas Chong, Jonathan Leung, Krish Shah, Alvin Yu, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) Votee AI Vector Institute(向量研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted to ACL 2025 main, camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11865 2025-05-19 cs.CV 83%

From Image to Video, what do we need in multimodal LLMs?

Suyuan Huang, Haoxin Zhang, Linqing Zhong, Honggu Chen, Yan Gao, Yao Hu, Zengchang Qin

机构 * Intelligent Computing and Machine Learning Lab, School of ASEE, Beihang University(北京航空航天大学自动化学院智能计算与机器学习实验室) Xiaohongshu(小红书) School of Sino-French Engineer, Beihang University(北京航空航天大学中法工程师学院) College of Engineering and Computer Science, VinUniversity(Vin大学工程与计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05714 2025-05-12 cs.CL 83%

TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries

Jinze Lv, Jian Chen, Zi Long, Xianghua Fu, Yin Chen

机构 * College of Application and Technology, Shenzhen University, China(应用技术学院,深圳大学,中国) College of Big Data and Internet, Shenzhen Technology University, China(大数据与互联网学院,深圳科技大学,中国)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments NLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02096 2025-05-06 cs.MM 83%

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing

Yaru Chen, Peiliang Zhang, Fei Li, Faegheh Sardari, Ruohao Guo, Zhenbo Li, Wenwu Wang

专题命中 视频多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.MM

Comments Accepted by ICMR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13983 2025-04-14 cs.CV 83%

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

Jiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li, Jiannan Ge, Hongtao Xie, Yongdong Zhang

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03735 2025-04-02 cs.CV 83%

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Chaoyu Li, Eun Woo Im, Pooyan Fazli

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23660 2025-04-01 cs.CV 83%

DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance

Junjie Zheng, Zihao Chen, Chaofan Ding, Xinhan Di

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19406 2025-03-26 cs.CV 83%

M$^2$CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation

Ziyuan Liu, Jiawei Zhang, Wenyu Wang, Yuantao Gu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19134 2025-03-26 cs.CL cs.CR 83%

MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks

Wenhao You, Bryan Hooi, Yiwei Wang, Youke Wang, Zong Ke, Ming-Hsuan Yang, Zi Huang, Yujun Cai

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13281 2025-03-25 cs.CV cs.AI cs.CL cs.MM 83%

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

Ziyang Luo, Haoning Wu, Dongxu Li, Jing Ma, Mohan Kankanhalli, Junnan Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments CVPR 2025, Project Page: https://videoautoarena.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10523 2025-03-14 cs.CV 83%

Interactive Multimodal Fusion with Temporal Modeling

Jun Yu, Yongqi Wang, Lei Wang, Yang Zheng, Shengfan Xu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17599 2025-03-14 cs.CL 83%

MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference

Zhongwei Wan, Hui Shen, Xin Wang, Che Liu, Zheda Mai, Mi Zhang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments NAACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20105 2024-12-31 cs.CV 83%

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

Jiedong Zhuang, Lu Lu, Ming Dai, Rui Hu, Jian Chen, Qiang Liu, Haoji Hu

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18060 2024-12-25 cs.CV 83%

An Ensemble Approach to Short-form Video Quality Assessment Using Multimodal LLM

Wen Wen, Yilin Wang, Neil Birkbeck, Balu Adsumilli

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14006 2024-12-19 cs.CV 83%

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Cong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng, Yong Liu, Zheng Zhao, Yujiu Yang

专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15829 2024-12-03 cs.CV 83%

SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization

Sicheng Liu, Lintao Wang, Xiaogang Zhu, Xuequan Lu, Zhiyong Wang, Kun Hu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures, submitted to ACM Multimedia Asia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17773 2024-12-02 cs.CV 83%

XTrack: Multimodal Training Boosts RGB-X Video Object Trackers

Yuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou, Guolei Sun, Eduard Zamfi, Chao Ma, Danda Pani Paudel, Luc Van Gool, Radu Timofte

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11pages, 5figs

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09875 2024-10-15 cs.CV cs.IR 83%

ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification

Chen Mao, Chong Tan, Jingqi Hu, Min Zheng

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04955 2024-10-01 cs.CV 83%

Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations

Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang, Peng Zhai, Song Wang, Lihua Zhang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by TCSVT 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05930 2024-09-18 cs.CV 83%

MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer

Rezaul Karim, He Zhao, Richard P. Wildes, Mennatullah Siam

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

Comments Extension of CVPR'23 paper for journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10213 2024-09-17 cs.CV 83%

Neuromorphic Facial Analysis with Cross-Modal Supervision

Federico Becattini, Luca Cultrera, Lorenzo Berlincioni, Claudio Ferrari, Andrea Leonardo, Alberto Del Bimbo

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted for publication at the ECCV 2024 workshop on Neuromorphic Vision: Advantages and Applications of Event Cameras (NEVI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09362 2024-09-17 cs.CL 83%

Generating Event-oriented Attribution for Movies via Two-Stage Prefix-Enhanced Multimodal LLM

Yuanjie Lyu, Tong Xu, Zihan Niu, Bo Peng, Jing Ke, Enhong Chen

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏