arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4772 篇

2506.23972 2025-10-02 cs.CV 83%

Learning Frequency and Memory-Aware Prompts for Multi-Modal Object Tracking

Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) School of Information Science and Engineering(信息科学与工程学院) School of Mathematics(数学学院) Purple Mountain Laboratories(紫金山实验室)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26473 2025-10-01 cs.AI 83%

STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models

Shaoxiong Guo, Tianyi Du, Lijun Li, Yuyao Wu, Jie Li, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) East China Normal University(华东师范大学) Soochow University(苏州大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23284 2025-09-30 cs.CV 83%

Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval

Dohwan Ko, Ji Soo Lee, Minhyuk Choi, Zihang Meng, Hyunwoo J. Kim

机构 * Korea University(韩国大学) Meta GenAI(Meta 生成人工智能) KAIST(韩国科学技术院)

专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21100 2025-09-26 cs.CV 83%

VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception

Ziang Yan, Xinhao Li, Yinan He, Zhengrong Yue, Xiangyu Zeng, Yali Wang, Yu Qiao, Limin Wang, Yi Wang

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室) Nanjing University(南京大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02889 2025-09-24 cs.CL cs.AI cs.CV cs.MM 83%

LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture

Xidong Wang, Dingjie Song, Shunian Chen, Junyin Chen, Zhenyang Cai, Chen Zhang, Lichao Sun, Benyou Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Lehigh University(莱斯利大学) Meituan(美团)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15178 2025-09-19 cs.CV 83%

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

Zaiquan Yang, Yuhao Liu, Gerhard Hancke, Rynson W. H. Lau

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Journal ref NeurIPS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14893 2025-09-19 cs.SD eess.AS 83%

Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic event Classification

Yuanjian Chen, Yang Xiao, Jinjie Huang

机构 * Harbin University of Science(哈尔滨理工大学) Technology The University of Melbourne(技术墨尔本大学)

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12893 2025-09-17 cs.CV 83%

MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization

Yiyi Zhang, Yuchen Yuan, Ying Zheng, Jialun Pei, Jinpeng Li, Zheng Li, Pheng-Ann Heng

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18407 2025-09-10 cs.GR cs.CV 83%

IntuiTF: MLLM-Guided Transfer Function Optimization for Direct Volume Rendering

Yiyao Wang, Bo Pan, Ke Wang, Han Liu, Jinyuan Mao, Yuxin Liu, Minfeng Zhu, Xiuqi Huang, Weifeng Chen, Bo Zhang, Wei Chen

机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) Laboratory of Art and Archaeology Image (Zhejiang University), Ministry of Education, China(浙江大学艺术与考古图像实验室(教育部,中国)) Zhejiang University of Finance&Economics(浙江财经大学) Zhejiang University(浙江大学)

专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19787 2025-09-09 cs.LG cs.AI 83%

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives

Armin Saghafian, Amirmohammad Izadi, Negin Hashemi Dijujin, Mahdieh Soleymani Baghshah

机构 * Sharif University of Technology(谢里夫理工大学)

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments Accepted to TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04416 2025-09-04 cs.CV 83%

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Haoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma, Sule Bai, Chubin Zhang, Bowen Zhang, Zhichao Zhou, Dongliang He, Yansong Tang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17163 2025-08-27 cs.CV 83%

PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion Recognition

Kai Cui, Jia Li, Yu Liu, Xuesong Zhang, Zhenzhen Hu, Meng Wang

机构 * School of Instrument Science and Opto-electronics Engineering, Hefei University of Technology(仪器科学与光电工程学院,合肥工业大学) School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments To appear in IEEE TCSS. The source code is publicly available at https://github.com/MSA-LMC/PhysioSync

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10572 2025-08-15 cs.CV 83%

Towards Agentic AI for Multimodal-Guided Video Object Segmentation

Tuyen Tran, Thao Minh Le, Truyen Tran

机构 * Applied Artificial Intelligence Institute, Deakin University, Australia(应用人工智能研究所,德金大学,澳大利亚)

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09789 2025-08-14 cs.IR cs.CV 83%

Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations

Marco De Nadai, Andreas Damianou, Mounia Lalmas

机构 * Spotify Denmark(Spotify丹麦分公司) Spotify United Kingdom(Spotify英国分公司)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16596 2025-08-05 cs.CV 83%

A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization

Wenbo Xu, Junyan Wu, Wei Lu, Xiangyang Luo, Qian Wang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(计算机科学与工程学院,中山大学) State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室) School of Cyber Science and Engineering, Wuhan University(网络科学与工程学院,武汉大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 9 pages, 3 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20988 2025-07-31 cs.CL cs.GR 83%

Cross-Modal State-Space Graph Reasoning for Structured Summarization

Hannah Kim, Sofia Martinez, Jason Lee

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21945 2025-07-30 cs.CV 83%

Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

Xin Wang, Peng-Jie Li, Yuan-Yuan Shen

机构 * School of Sport Engineering, Beijing Sport University(体育工程学院,北京体育大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to Applied Soft Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15569 2025-07-22 cs.CV 83%

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

Xiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng, Xiaofeng Wang, Yun Zheng, Xingang Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团) Peking University(北京大学) Luoyang Institute for Robot and Intelligent Equipment(洛阳机器人与智能装备研究所)

专题命中 视频多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16450 2025-06-23 cs.CV 83%

How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?

Giuseppe Lando, Rosario Forte, Giovanni Maria Farinella, Antonino Furnari

机构 * Department of Mathematics and Computer Science(数学与计算机科学系)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10430 2025-06-13 cs.CV 83%

MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment

Shuo wang, Jihao Zhang

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09668 2025-06-12 cs.CV cs.LG 83%

CINeMA: Conditional Implicit Neural Multi-Modal Atlas for a Spatio-Temporal Representation of the Perinatal Brain

Maik Dannecker, Vasiliki Sideri-Lampretsa, Sophie Starck, Angeline Mihailov, Mathieu Milh, Nadine Girard, Guillaume Auzias, Daniel Rueckert

机构 * School of Computation, Information and Technology, and the School of Medicine and Health(计算信息学院及医学健康学院) Technical University of Munich(慕尼黑技术大学) Department of Computing(计算系) Institut de Neurosciences de la Timone, UMR 7289, CNRS, Aix-Marseille Université(神经科学研究所,UMR 7289,CNRS,艾克斯-马赛大学) Aix-Marseille Univ, APHM, Service de Neuroradiologie Diagnostique et Interventionnelle, Hôpital de la Timone(艾克斯-马赛大学,APHM,诊断与介入神经放射科,泰米翁医院) Aix-Marseille Univ, APHM, service de neurologie pédiatrique, Hôpital de la Timone(艾克斯-马赛大学,APHM,儿童神经科,泰米翁医院)

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Work currently under revision for IEEE TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08649 2025-06-11 cs.CV 83%

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization

Zhiyi Zhu, Xiaoyu Wu, Youwei Lu

机构 * Department of Information and Communication Engineering(信息与通信工程系)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08466 2025-06-10 cs.CV 83%

Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

Quan Zhang, Jinwei Fang, Rui Yuan, Xi Tang, Yuxin Qi, Ke Zhang, Chun Yuan

机构 * Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to CVPR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00304 2025-06-04 cs.CV 83%

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

Yuxiang Guo, Faizan Siddiqui, Yang Zhao, Rama Chellappa, Shao-Yuan Lo

机构 * Johns Hopkins University(约翰霍普金斯大学) Honda Research Institute USA(本田研究院美国)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Paper is accepted by IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08282 2025-06-03 cs.CV 83%

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Hongyu Li, Jinyu Chen, Ziyu Wei, Shaofei Huang, Tianrui Hui, Jialin Gao, Xiaoming Wei, Si Liu

机构 * School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Meituan(美团)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24476 2025-06-02 cs.CV 83%

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model

Yuting Zhang, Hao Lu, Qingyong Hu, Yin Wang, Kaishen Yuan, Xin Liu, Kaishun Wu

机构 * The Hong Kong University of Science & Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science & Technology(香港科技大学) Zhejiang University(浙江大学) Lappeenranta-Lahti University of Technology(拉佩兰塔-拉赫蒂技术大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19400 2025-05-27 cs.AI cs.CL cs.CV cs.MM 83%

TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding

Max Ku, Thomas Chong, Jonathan Leung, Krish Shah, Alvin Yu, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) Votee AI Vector Institute(向量研究所)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted to ACL 2025 main, camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11865 2025-05-19 cs.CV 83%

From Image to Video, what do we need in multimodal LLMs?

Suyuan Huang, Haoxin Zhang, Linqing Zhong, Honggu Chen, Yan Gao, Yao Hu, Zengchang Qin

机构 * Intelligent Computing and Machine Learning Lab, School of ASEE, Beihang University(北京航空航天大学自动化学院智能计算与机器学习实验室) Xiaohongshu(小红书) School of Sino-French Engineer, Beihang University(北京航空航天大学中法工程师学院) College of Engineering and Computer Science, VinUniversity(Vin大学工程与计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05714 2025-05-12 cs.CL 83%

TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries

Jinze Lv, Jian Chen, Zi Long, Xianghua Fu, Yin Chen

机构 * College of Application and Technology, Shenzhen University, China(应用技术学院,深圳大学,中国) College of Big Data and Internet, Shenzhen Technology University, China(大数据与互联网学院,深圳科技大学,中国)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments NLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02096 2025-05-06 cs.MM 83%

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing

Yaru Chen, Peiliang Zhang, Fei Li, Faegheh Sardari, Ruohao Guo, Zhenbo Li, Wenwu Wang

专题命中 视频多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.MM

Comments Accepted by ICMR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏