arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2605.15071 2026-05-15 cs.CV cs.AI cs.CL 67%

On the Cultural Anachronism and Temporal Reasoning in Vision Language Models

关于视觉语言模型中文化错位与时间推理的问题

Mukul Ranjan, Prince Jha, Khushboo Kumari, Zhiqiang Shen

机构 * MBZUAI Inception

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文探讨了视觉语言模型在解读历史文物时存在的文化错位问题,通过设计TAB-VLM基准测试集评估模型的时间推理能力,发现现有模型在处理非西方文化文物时存在显著缺陷。

Comments Project Page: https://khushboo0012.github.io/tab-vlm-webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01707 2026-05-14 cs.CV cs.AI cs.CL 67%

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

StreamGaze:流媒体视频中的目光引导时间推理与前瞻性理解

Daeun Lee, Subhojyoti Mukherjee, Branislav Kveton, Ryan A. Rossi, Viet Dac Lai, Seunghyun Yoon, Trung Bui, Franck Dernoncourt, Mohit Bansal

机构 * University of North Carolina, Chapel Hill(北卡罗来纳大学教堂山分校) Adobe Research(Adobe研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 StreamGaze通过引入目光引导的任务评估多模态大语言模型在流媒体视频中的时间推理和前瞻性理解能力,揭示了现有模型在目光引导下的时间推理、意图建模和前瞻性预测方面的局限性。

Comments Accepted to CVPR 2026 with strong scores (5/5/5) but desk-rejected after the camera-ready due to not completing all reviewing duties

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14724 2026-05-08 cs.CV cs.AI cs.CL 67%

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

HERMES: KV缓存作为分层内存用于高效视频流理解

Haowei Zhang, Shudong Yang, Jinlan Fu, See-Kiong Ng, Xipeng Qiu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 HERMES提出一种无需训练的架构,通过将KV缓存视为分层内存框架,实现视频流的实时准确理解,提升处理速度并降低内存消耗。

Comments Accepted to ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28007 2026-05-01 q-bio.NC 67%

Multisensory learning recruits visual neurons into an olfactory memory engram

多感官学习招募视觉神经元进入嗅觉记忆表征

Zeynep Okray, Nils Otto, Anna A. Cook, Clifford Talbot, Ashwin Miriyala, Martín Klappenbach, Ciara Stern, Kieran Desmond, Paola Vargas-Gutierrez, Scott Waddell

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 研究揭示多感官学习通过扩大记忆表征提升记忆表现,核心方法涉及视觉神经元与嗅觉记忆的连接,主要贡献是发现多感官训练增强记忆表达的机制。

Comments 24 pages, 9 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07180 2026-05-01 cs.CL cs.AI cs.CV 67%

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs

视频中的奉承:视频大语言模型中奉承行为的基准测试与分析

Wenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang, Zikun Guo, Yuyu Luo, Lijie Hu, Di Wang

机构 * Provable Responsible AI and Data Analytics (PRADA) Lab(可证负责任人工智能与数据 analytics 实验室) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹科学与技术大学) HKUST(香港科技大学) MBZUAI(穆罕默德·本·拉希德人工智能研究所) University of Science and Technology of China(中国科学技术大学) Kyungpook National University(庆尚国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出VISE基准,用于评估视频大语言模型在面对误导性输入时的奉承行为,通过多类型分析和两种无训练缓解策略提升模型可靠性。

Comments 27 Pages, Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24329 2026-04-14 cs.CL cs.AI cs.CV 67%

GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents

GameplayQA: 一种用于3D虚拟代理多视频理解的决策密集型视角同步评估框架

Yunzhe Wang, Runhui Xu, Kexin Zheng, Tianyi Zhang, Jayavibhav Niranjan Kogundi, Soham Hans, Volkan Ustun

机构 * University of Southern California(南加州大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 GameplayQA通过密集标注多玩家3D游戏视频,构建了2.4K诊断QA对,评估代理感知与推理能力,揭示了前沿MLLM在时间同步、跨视频定位及决策密度处理上的不足。

Comments Accepted to the Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05117 2026-04-08 cs.CV cs.AI cs.CL 67%

Watch Before You Answer: Learning from Visually Grounded Post-Training

在回答前观看:从视觉引导的后训练中学习

Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang, Penghui Du, Yiming Jia, Dongfu Jiang, Xuan He, Shenhui Zhang, Ping Nie, Peter West, Kelsey R. Allen

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Etude AI Kolors Team, Kuaishou Technology(快手科技Kolors团队) University of Toronto(多伦多大学) University of Waterloo(滑铁卢大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文发现现有视频理解基准中40-60%的问题可通过文本线索回答,提出VidGround方法通过仅使用视觉引导问题提升VLM性能,实验显示其在后训练中效果优于复杂方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04953 2026-04-08 cs.CV cs.AI cs.HC cs.IR cs.MM 67%

Generative AI for Video Trailer Synthesis: From Extractive Heuristics to Autoregressive Creativity

生成AI用于视频预告片合成:从提取启发式方法到自回归创造力

Abhishek Dharmaratnakar, Srivaths Ranganathan, Debanshu Das, Anushree Sinha

机构 * Google LLC(谷歌有限责任公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文综述了自动视频预告片生成领域从启发式提取到深度生成合成的转变,探讨了自回归Transformer、LLM协同流程和文本到视频基础模型等生成技术,并讨论了高保真神经合成的伦理挑战。

Comments 7 pages, 3 figures, accepted in WSDM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04379 2026-01-30 cs.CV cs.AI cs.CL cs.HC 67%

Can Large Language Models Capture Video Game Engagement?

大语言模型能否捕捉视频游戏参与度?

David Melhart, Matthew Barthet, Georgios N. Yannakakis

机构 * Institute of Digital Games, University of Malta Msida, Malta(马耳他大学数字游戏研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文研究了大语言模型在多模态输入下预测视频游戏参与度的能力,发现尽管LLMs在多个领域表现优异,但其在连续情绪标注上仍无法超越人类注释。

Comments This work has been submitted to the IEEE for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22226 2025-12-30 cs.CV cs.AI cs.CL cs.LG 67%

VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs

VideoScaffold: 基于弹性尺度的视觉层次结构用于多模态大语言模型中的流媒体视频理解

Naishan Zheng, Jie Huang, Qingpei Guo, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Ant Group(蚂蚁集团)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 VideoScaffold通过弹性尺度事件分割和层次事件整合,实现了流媒体视频理解中的细粒度到抽象事件推理的动态转换,提升多模态大语言模型的视频处理性能。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12284 2025-12-25 eess.IV cs.AI cs.AR cs.CV cs.MM 67%

V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache Retrieval

V-Rex: 通过动态KV缓存检索实现实时视频大语言模型加速

Donghyuk Kim, Sejeong Yang, Wonjin Shin, Joo-Young Kim

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 V-Rex通过动态KV缓存检索算法和硬件加速器,实现边缘设备上的实时视频LLM推理,显著提升速度和能效。

Comments 14 pages, 20 figures, conference, accepted by HPCA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18318 2025-12-23 cs.MM cs.AI cs.CV cs.DC cs.NI 67%

Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems

异步流水线并行用于实时多语言唇同步视频通信系统

Eren Caglar, Amirkia Rafiei Oskooei, Mehmet Kutanoglu, Mustafa Keles, Mehmet S. Aktas

机构 * Department of Data Science Big Data Yildiz Technical University Istanbul, Turkey Department of Computer Engineering Yildiz Technical University Istanbul, Turkey R\&D Center Aktif Investment Bank Inc. Istanbul, Turkey 3.5cm R\&D Center 3.5cm Aktif Investment Bank Inc. 3.5cm Istanbul, Turkey 3.5cm 0.25cm Department of Computer Engineering 0.25cm Yildiz Technical University 0.25cm Istanbul, Turkey 0.25cm

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出一种异步流水线并行的Transformer框架,用于实时多语言唇同步,通过模块化设计和优化技术提升处理速度与资源利用率。

Comments Accepted to IEEE Big Data 2025, AIDE4IoT Workshop. Copyright \c{opyright} 2025 IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08051 2025-12-12 cs.LG 67%

ST-GraphNet: A Spatio-Temporal Graph Neural Network for Understanding and Predicting Automated Vehicle Crash Severity

ST-GraphNet:一种用于理解和预测自动驾驶车辆碰撞严重程度的时空图神经网络

Mahmuda Sultana Mimi, Md Monzurul Islam, Anannya Ghosh Tusti, Shriyank Somvanshi, Subasish Das

机构 * Texas State University(德克萨斯州立大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract)

AI总结 ST-GraphNet通过时空图神经网络整合多模态数据,以高准确率预测自动驾驶车辆碰撞严重程度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09841 2025-12-11 cs.CL cs.CV cs.MM 67%

ChronusOmni: Improving Time Awareness of Omni Large Language Models

ChronusOmni: 提升 Omni 大语言模型的时间感知能力

Yijing Chen, Yihan Wu, Kaisi Guan, Yuchen Ren, Yuyue Wang, Ruihua Song, Liyun Ru

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学光荣人工智能学院) Baichuan Inc.(百川科技公司)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 ChronusOmni 通过增强时间感知能力,提升音频视觉时间定位任务的性能,实现跨模态统一建模和细粒度推理。

Comments Code available at https://github.com/YJCX330/Chronus/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23400 2025-12-05 astro-ph.IM astro-ph.SR 67%

Solar flare forecasting with foundational transformer models across image, video, and time-series modalities

基于基础Transformer模型的太阳耀斑预测:跨图像、视频和时间序列模态

S. Riggi, P. Romano, A. Pilzer, U. Becciani

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 本文通过比较三种基础Transformer模型在太阳耀斑预测中的表现,发现Moirai2在时间序列预测中表现最佳,展示了预训练模型在多模态空间天气预测中的潜力。

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16595 2025-11-27 cs.CV cs.AI cs.CL 67%

TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding

TimeViper: 一种用于高效长视频理解的混合Mamba-Transformer视觉语言模型

Boshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, Qin Jin

机构 * AIM3 Lab, Renmin University of China(中国人民大学人工智能实验室) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 TimeViper是一种混合Mamba-Transformer模型,通过TransV模块实现高效长视频理解,提升多模态处理能力。

Comments Project page: https://xuboshen.github.io/TimeViper; Code: https://github.com/xiaomi-research/timeviper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19475 2025-11-26 cs.CV cs.AI cs.MM 67%

Tracking and Segmenting Anything in Any Modality

任何模态下任何事物的跟踪与分割

Tianlu Zhang, Qiang Zhang, Guiguang Ding, Jungong Han

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 SATA提出了一种通用框架,通过解耦的专家混合机制和任务感知多目标跟踪流程,统一了多种跟踪和分割任务,提升了模型的泛化能力。

Comments Accpetd by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19002 2025-11-18 cs.CV cs.AI cs.CL cs.LG 67%

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

Hao Wang, Eiki Murata, Lingfang Zhang, Ayako Sato, So Fukuda, Ziqi Yin, Wentao Hu, Keisuke Nakao, Yusuke Nakamura, Sebastian Zwirner, Yi-Chia Chen, Hiroyuki Otomo, Hiroki Ouchi, Daisuke Kawahara

机构 * Waseda University(早稻田大学) CyberAgent, Inc.(CyberAgent公司) AI Shift, Inc.(AI Shift公司) Nara Institute of Science and Technology(奈良研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10008 2025-11-18 cs.MM cs.AI cs.CV 67%

Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Updated with the ICIDS 2025 camera-ready version. This revision includes the final title, updated abstract, improved explanations of the narrative coherence framework, and minor editorial changes. Figures and examples have been refined for clarity. No new experiments were added

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16781 2025-11-14 cs.CV cs.AI cs.CL 67%

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

Shihao Ji, Zihui Song

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper is being withdrawn because we have identified a significant error in the implementation of our self-supervised clustering approach. Specifically, our feature aggregation step inadvertently leaked temporal information across frames, which violates the core assumption of our training-free method. We sincerely apologize to the research community

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21786 2025-10-28 cs.CV cs.AI cs.MM 67%

EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction

Qile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang, Chao Tong

机构 * Beihang University(北京航空航天大学) University of Science and Technology Beijing(北京科技大学) School of Computer Science and Engineering(计算机科学与工程学院) State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 15 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12299 2025-10-15 cs.IR 67%

An Empirical Study for Representations of Videos in Video Question Answering via MLLMs

Zhi Li, Yanan Wang, Hao Niu, Julio Vizcarra, Masato Taya

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract)

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25745 2025-10-01 cs.CV cs.CL cs.MM 67%

FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos

Siddhant Sukhani, Yash Bhardwaj, Riya Bhadani, Veer Kejriwal, Michael Galarnyk, Sudheer Chava

机构 * Stanford University(斯坦福大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments ICCV Short Video Understanding Workshop Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07966 2025-10-01 cs.CV cs.AI cs.CL 67%

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu, Hanrong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, Sifei Liu, Hongxu Yin, Yao Lu, Song Han

机构 * NVIDIA MIT(麻省理工学院) HKU(香港大学) UC Berkeley(加州大学伯克利分校)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS 2025. Code at https://github.com/NVlabs/Long-RL and model at https://huggingface.co/Efficient-Large-Model/LongVILA-R1-7B

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11915 2025-09-18 cs.SD cs.CV cs.LG cs.MM eess.AS 67%

Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound

Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam

机构 * Graduate School of AI, KAIST(韩国国立庆熙大学人工智能研究生院) Graduate School of CT, KAIST(韩国国立庆熙大学CT研究生院)

专题命中 视频多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12231 2025-09-17 cs.DC 67%

Research on fault diagnosis and root cause analysis based on full stack observability

Jian Hou

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04957 2025-09-08 cs.CV cs.MM cs.SD eess.AS 67%

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

Gehui Chen, Guan'an Wang, Xiaowen Huang, Jitao Sang

机构 * School of Computer Science Technology, Beijing Jiaotong University Beijing China Beijing Key Laboratory of Traffic Data Mining Key Laboratory of Big Data \& Artificial Intelligence in Transportation, Ministry of Education Beijing China Technology, Beijing Jiaotong University Key Laboratory of Big Data \& Artificial Intelligence in Transportation, Ministry of Education

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13919 2025-09-03 cs.CV cs.AI cs.CL cs.LG cs.RO 67%

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang, Orr Zohar, Zeyu Wang, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21586 2025-08-08 cs.CL cs.AI cs.CV 67%

Can Vision Language Models Understand Mimed Actions?

Hyundong Cho, Spencer Lin, Tejas Srinivasan, Michael Saxon, Deuksin Kwon, Natali T. Chavez, Jonathan May

机构 * Information Sciences Institute(信息科学研究所) Institute for Creative Technologies(创意技术研究所) Department of Computer Science(计算机科学系) University of Southern California(南加州大学) University of California, Santa Barbara(加州大学圣巴巴拉分校) Aristotle University of Thessaloniki(希腊雅典纳大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09068 2025-07-24 cs.CV cs.AI cs.IR cs.LG cs.MM 67%

Infinite Video Understanding

Dell Zhang, Xiangyu Chen, Jixiang Luo, Mengxi Jia, Changzhi Sun, Ruilong Ren, Jingren Liu, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) Peking University(北京大学) Tianjin University(天津大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏