arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4772 篇

2408.12322 2026-07-13 cs.CV 版本更新 88%

Zero-shot 3D General Obstacle Detection via Multimodal Foundation Models and Geometry

通过多模态基础模型和几何进行零样本3D通用障碍物检测

Tamás Matuszka, Péter Hajas, Dávid Szeghy

机构 * aiMotive

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

AI总结 针对自动驾驶中通用障碍物检测难题,提出结合多模态基础模型与几何推理的免训练零样本方法,能精准定位达100米,借基础模型先验提升召回率,还可实现可扩展自动标注。

Comments Accepted to CVPR 2026 AUTOPILOT Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29425 2026-05-29 cs.AI 88%

ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control

ReasonLight: 一种多模态基础模型增强的强化学习框架用于零样本交通信号控制

Aoyu Pang, Maonan Wang, Yuejiao Xie, Chung Shue Chen, Zhiwei Yang, Man-On Pun

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hong Kong(香港中文大学机械与自动化工程系) Shanghai AI Laboratory, Shanghai, China(上海人工智能实验室) Nokia Bell Labs, Paris-Saclay, France(法国巴黎萨克雷诺基贝尔实验室)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.AI

AI总结 提出ReasonLight框架,通过多模态基础模型增强强化学习,利用路侧传感器和摄像头数据实现零样本适应罕见交通事件,显著降低紧急车辆等待时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00209 2026-02-03 cs.MM 88%

Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization

分而治之:通过跨模态融合与定位实现多模态视频深度伪造检测

Qingcao Li, Miao He, Liang Yi, Qing Wen, Yitao Zhang, Hongshuo Jin, Peng Cheng, Zhongjie Ba, Li Lu, Kui Ren

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(title);audio-visual(abstract);分类 cs.MM

AI总结 本文提出一种通过跨模态融合与定位实现多模态视频深度伪造检测的系统,通过音频和视觉模块的融合提升检测鲁棒性。

Comments The 3rd Place, IJCAI 2025 Workshop on Deepfake Detection, Localization, and Interpretability

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00696 2025-11-10 cs.LG cs.AI 88%

CTPD: Cross-Modal Temporal Pattern Discovery for Enhanced Multimodal Electronic Health Records Analysis

Fuying Wang, Feng Wu, Yihan Tang, Lequan Yu

机构 * Department of Statistics and Actuarial Science, School of Computing and Data Science, The University of Hong Kong(统计与精算学系,计算与数据科学学院,香港大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19406 2024-12-30 cs.CV 88%

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios

Jiaqi Fan, Jianhua Wu, Jincheng Gao, Jianhao Yu, Yafei Wang, Hongqing Chu, Bingzhao Gao

机构 * Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统研究院) School of Automotive Studies, Tongji University(同济大学汽车学院) School of Mechanical Engineering, Shanghai Jiao Tong University(上海交通大学机械与动力工程学院)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12628 2024-12-19 cs.CV 88%

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

Ziheng Zhou, Jinxing Zhou, Wei Qian, Shengeng Tang, Xiaojun Chang, Dan Guo

专题命中 视频多模态 :cross-modal(title,abstract);audio-visual(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2025. Project page: https://github.com/zzhhfut/CCNet-AAAI2025. Jinxing Zhou and Dan Guo are the corresponding authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00103 2024-12-03 cs.RO cs.AI cs.LG 88%

MLLM-Search: A Zero-Shot Approach to Finding People using Multimodal Large Language Models

Angus Fung, Aaron Hao Tan, Haitong Wang, Beno Benhabib, Goldie Nejat

专题命中 视频多模态 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12374 2022-07-13 cs.CV 88%

MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video Parsing

Jiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng, Yuejie Zhang

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV

Comments ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.03501 2020-06-09 cs.CV cs.LG 88%

Cross-modal Learning for Multi-modal Video Categorization

Palash Goyal, Saurabh Sahu, Shalini Ghosh, Chul Lee

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18552 2025-07-25 cs.CV cs.AI 88%

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding

Baoyao Yang, Wanyun Li, Dixin Chen, Junxiang Chen, Wenbin Yao, Haifeng Lin

机构 * Guangdong University of Technology(广东工业大学) Wechat, Tencent(微信、腾讯)

专题命中 视频多模态 :omni-modal(title,abstract);multi-modal(abstract);MLLM(abstract);cross-modal(abstract)

Comments 7 pages; 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07433 2026-06-08 cs.CV cs.AI cs.MM 新提交 88%

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

Watch, Remember, Reason: 基于多模态大语言模型的人类视角视频理解

Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang

机构 * School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院) Wuhan University(武汉大学) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) CASIA(中国科学院自动化研究所) University of Tokyo(东京大学) University of Liverpool(利物浦大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) UC Merced(加州大学默塞德分校)

专题命中 视频多模态 :MLLM(summary_cn,abstract);multimodal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 提出人类视角下视频理解的三个功能能力(观看、记忆、推理),构建统一框架分析视频MLLM的感知、记忆、推理和预测,并总结挑战、方法、应用及未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25678 2026-03-03 cs.LG 88%

Massively Multimodal Foundation Models: A Framework for Capturing Interactions with Specialized Mixture-of-Experts

大规模多模态基础模型:一种捕捉交互的专用专家混合框架

Xing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang, Suchi Saria

机构 * Johns Hopkins University(约翰霍普金斯大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title);cross-modal(abstract)

AI总结 本文提出了一种大规模多模态基础模型框架,通过量化模态间的时间依赖性,改进混合专家路由机制,提升跨模态交互处理能力。

Comments Published at International Conference on Learning Representations (ICLR) 2026 as a conference paper. 28 pages, 16 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04073 2026-01-08 cs.CV cs.AI cs.CL 88%

Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts

在跨模态冲突下分析大型多模态模型的推理一致性

Zhihao Zhu, Jiafeng Liang, Shixin Jiang, Jinlan Fu, Ming Liu, Guanglu Sun, See-Kiong Ng, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Peng Cheng Laboratory(鹏城实验室) National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(title);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出逻辑图扰动协议,通过主动视觉上下文细化方法,解决大型多模态模型在跨模态冲突下的推理一致性问题,提升推理鲁棒性。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06334 2025-12-09 cs.IR 88%

Enhanced Multimodal Video Retrieval System: Integrating Query Expansion and Cross-modal Temporal Event Retrieval

增强多模态视频检索系统:整合查询扩展与跨模态时间事件检索

Van-Thinh Vo, Minh-Khoi Nguyen, Minh-Huy Tran, Anh-Quan Nguyen-Tran, Duy-Tan Nguyen, Khanh-Loi Nguyen, Anh-Minh Phan

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(title,abstract)

AI总结 本文提出了一种增强多模态视频检索系统,通过整合查询扩展和跨模态时间事件检索,提升检索精度和效率。

Comments 11 pages, 6 figures, SOICT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17084 2026-08-19 cs.CL 新提交 87%

Uncertainty-Aware Decision Making in Multimodal Large Language Models

多模态大语言模型中的不确定性感知决策

Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed

机构 * Khalifa University of Science and Technology(哈利法科技大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CL

AI总结 本调查围绕决策中心框架,综述多模态大语言模型(MLLMs)的不确定性感知决策相关研究,对比同类调查并指出源感知分解等开放问题,核心是不确定性需改善多模态证据下的系统行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22897 2026-07-03 cs.AI cs.CL cs.CV cs.LG cs.MM 版本更新 87%

OmniGAIA: Towards Native Omni-Modal AI Agents

OmniGAIA:迈向原生全模态AI代理

Xiaoxi Li, Wenxiang Jiao, Jiarui Jin, Haoxuan Li, Hao Wang, Shijian Wang, Guanting Dong, Jiajie Jin, Yinuo Wang, Yuan Lu, Ji-Rong Wen, Zhicheng Dou, Zhouchen Lin

机构 * Renmin University of China(中国人民大学) Xiaohongshu Inc.(小红书公司) Peking University(北京大学) Southeast University(东南大学) Tsinghua University(清华大学)

专题命中 视频多模态 :omni-modal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出OmniGAIA基准和OmniAtlas代理,通过全模态事件图和后见引导树探索策略,实现跨视频、音频和图像的深度推理与工具使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16495 2026-03-18 cs.AI 87%

ExpressMind: A Multimodal Pretrained Large Language Model for Expressway Operation

ExpressMind:一种用于高速公路运营的多模态预训练大语言模型

Zihe Wang, Yihuan Wang, Haiyang Yu. Zhiyong Cui, Xiaojian Liao, Chengcheng Wang, Yonglin Tian, Yongxin Tong

机构 * Beihang University(北航) Shandong Hi-speed Group Co., Ltd(山东高速集团有限公司) Institute of automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);cross-modal(abstract)

AI总结 本文提出ExpressMind,一种针对高速公路运营的多模态预训练大语言模型,通过构建全栈数据集和双层预训练方法,提升事件检测、安全响应生成和复杂交通分析能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02963 2026-07-07 cs.CV cs.AI cs.MM 新提交 87%

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

用于全模态密集视频字幕的并行自回归解码

Wenzheng Zeng, Siyi Jiao, Chen Gao, Hwee Tou Ng, Mike Zheng Shou

机构 * National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :omni-modal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 研究密集视频字幕生成,提出并行自回归框架,利用事件间弱局部依赖重组因果依赖图,引入全局规划和事件分解并行解码机制,提升效率与字幕性能。

Comments ECCV 2026. Project website: https://github.com/showlab/PadCaptioner

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07541 2026-06-09 cs.HC cs.AI cs.CV cs.CY cs.MM 新提交 87%

Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation

多模态大语言模型作为视频研究中的合成参与者:一项评估

Prabal Shrestha, Bohan Jiang, Haoning Xue, Huan Liu, Xinyi Zhou

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI、cs.MM

AI总结 本研究评估多模态大语言模型在视频感知任务中模拟人类主观评分的表现,发现模型存在偏差且与人类一致性有限。

Comments Accepted to SocialLLM @ ICWSM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19178 2025-09-30 cs.CV cs.AI cs.CL cs.IR cs.LG 87%

Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval

Yang Du, Yuqi Liu, Qin Jin

机构 * Renmin University of China(中国人民大学)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACMMM 2024 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19772 2025-03-21 cs.CV cs.CL cs.LG cs.MM 87%

LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Tiantian Geng, Jinrui Zhang, Qingni Wang, Teng Wang, Jinming Duan, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) University of Birmingham(伯明翰大学) University of Electronic Science and Technology of China(电子科技大学) The University of Hong Kong(香港大学) University of Manchester(曼彻斯特大学)

专题命中 视频多模态 :omni-modal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06020 2025-02-11 cs.CV cs.MM cs.SD eess.AS 87%

Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding

Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, Jiang Gui

机构 * Dartmouth College(达特茅斯学院)

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14200 2024-06-03 cs.CV cs.AI 87%

Awesome Multi-modal Object Tracking

Chunhui Zhang, Li Liu, Hao Wen, Xi Zhou, Yanfeng Wang

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);vision-language-audio(abstract);分类 cs.CV、cs.AI

Comments A continuously updated project to track the latest progress in multi-modal object tracking

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13024 2026-06-12 cs.LG cs.AI 新提交 86%

CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts

CausalMoE:基于模式路由异构专家的十亿规模多模态基础模型用于格兰杰因果发现

Bo Liu, Di Dai, Jingwei Liu, Jiarui Jin, Xiaocheng Fang, Guangkun Nie, Hongyan Li, Shenda Hong

机构 * State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室) National Institute of Health Data Science, and Institute for Artificial Intelligence, Peking University(北京大学健康医疗大数据国家研究院、人工智能研究院)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.AI

AI总结 提出CausalMoE,一种十亿规模多模态格兰杰因果基础模型,通过模式路由混合异构专家解耦动态机制,结合因果自注意力与LLM/VLM先验,实现稀疏因果图恢复,在监督和少样本场景中达到最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07642 2026-05-11 cs.CV 86%

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

EggHand:一种用于第一人称手姿态预测的多模态基础模型

Jaeyoung Choi, Hyeondong Kim, Yujin Kim, Daehee Park

机构 * DGIST, Republic of Korea(韩国忠南科学技术院)

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV

AI总结 本文提出EggHand模型,通过结合视觉-语言-动作模型的动作解码器和第一人称视频-文本编码器,实现对第一人称视频中手姿态序列的预测,提升了在剧烈视角变化下的鲁棒性和可控性。

Comments CVPR Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07497 2025-06-23 cs.CV 86%

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Xiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu, Lijun Zhou, Gangwei Xu, Shaoqing Xu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米电动车)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17088 2024-06-24 cs.CV 86%

Unsupervised Multimodal Deepfake Detection Using Intra- and Cross-Modal Inconsistencies

Mulin Tian, Mahyar Khayatkhoei, Joe Mathai, Wael AbdAlmageed

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(title);分类 cs.CV

Comments 11 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21022 2026-08-24 cs.CV cs.MM 新提交 86%

Recognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding

识别条件推理:一种用于细粒度微动作理解的无训练多模态大语言模型流水线

Fengshun Wang, Jin'ang Han, Zhigang Tu

机构 * Wuhan University(武汉大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.MM

AI总结 该研究提出一种无训练的多模态大语言模型流水线,动态分配子任务至适配的模型,在2026年MAC微动作挑战赛MA-Bench赛道的开放式任务上获显著性能优势,取得第一名。

Comments Accept at ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02927 2026-08-11 cs.CV cs.AI 86%

PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding

PrismVAU: 用于多模态视频异常理解的提示优化推理系统

Iñaki Erregue, Kamal Nasrollahi, Sergio Escalera

机构 * Universitat de Barcelona(巴塞罗那大学) Computer Vision Center(计算机视觉中心) Aalborg University(奥胡斯大学) Milestone Systems(Milestone系统)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 PrismVAU通过轻量级系统和自动提示工程实现高效的多模态视频异常理解,无需复杂标注和外部模块。

Comments This paper has been accepted to the 6th Workshop on Real-World Surveillance: Applications and Challenges (WACV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05833 2026-06-19 cs.CV cs.AI 版本更新 86%

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

从视频中学习几何表示以实现空间智能多模态大语言模型

Haibo Wang, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出GeoVR框架,通过从2D视频序列中蒸馏3D几何知识(包括相机姿态、深度图、尺度因子和多尺度3D特征),重塑多模态大语言模型的内部表示以赋予其空间智能,在空间推理基准上达到最先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏