arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-18 至 2025-11-18 共收录 142 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 22 篇

2511.13655 2025-11-18 cs.CV cs.LG 79%

OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation

Henry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng, Joseph Redmon, Hadrien Sablon, Ryan Park, Jacob Morrison, Alexandra Buraczynski, Karen Farley, Joshua Hansen, Andrew Howe, Patrick Alan Johnson, Mark Otterlee, Ted Schmitt, Hunter Pitelka, Stephen Daspit, Rachel Ratner, Christopher Wilhelm, Sebastian Wood, Mike Jacobi, Hannah Kerner, Evan Shelhamer, Ali Farhadi, Ranjay Krishna, Patrick Beukema

机构 * Allen Institute for AI(人工智能研究所) University of Washington(华盛顿大学) Arizona State University(亚利桑那州立大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12196 2025-11-18 cs.CV cs.HC 79%

Cross-View Cross-Modal Unsupervised Domain Adaptation for Driver Monitoring System

Aditi Bhalla, Christian Hellert, Enkelejda Kasneci

机构 * School of Social Sciences and Technology(社会科学与技术学院)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12027 2025-11-18 cs.CV cs.AI 73%

GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory

Jeong Hun Yeo, Sangyun Chung, Sungjune Park, Dae Hoe Kim, Jinyoung Moon, Yong Man Ro

机构 * Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST)) Visual Intelligence Research Section, Superintelligence Creative Research Laboratory, Electronics and Telecommunications Research Institute (ETRI)(视觉智能研究部,超智能创意研究实验室,电子电信研究院)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11908 2025-11-18 cs.CV cs.AI 73%

PI-NAIM: Path-Integrated Neural Adaptive Imputation Model

Afifa Khaled, Ebrahim Hamid Sumiea

机构 * University of Science and Technology of China(中国科学技术大学) Universiti Teknologi PETRONAS(Petronas科技大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13054 2025-11-18 cs.CV 70%

ViSS-R1: Self-Supervised Reinforcement Video Reasoning

Bo Fang, Yuxin Song, Qiangqiang Wu, Haoyuan Sun, Wenhao Wu, Antoni B. Chan

机构 * City University of Hong Kong(香港城市大学) Baidu Inc.(百度公司) Tsinghua University(清华大学) The University of Sydney(悉尼大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Our paper was initially titled "Video-SSR1: Self-Supervised Reinforcement Video Reasoning." Upon noticing its close resemblance to the title of a recently released paper, we have decided to rename our work as "ViSS-R1."

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11576 2025-11-18 cs.CV 70%

Causality Matters: How Temporal Information Emerges in Video Language Models

Yumeng Shi, Quanyu Long, Yin Wu, Wenya Wang

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19002 2025-11-18 cs.CV cs.AI cs.CL cs.LG 67%

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

Hao Wang, Eiki Murata, Lingfang Zhang, Ayako Sato, So Fukuda, Ziqi Yin, Wentao Hu, Keisuke Nakao, Yusuke Nakamura, Sebastian Zwirner, Yi-Chia Chen, Hiroyuki Otomo, Hiroki Ouchi, Daisuke Kawahara

机构 * Waseda University(早稻田大学) CyberAgent, Inc.(CyberAgent公司) AI Shift, Inc.(AI Shift公司) Nara Institute of Science and Technology(奈良研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10008 2025-11-18 cs.MM cs.AI cs.CV 67%

Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Updated with the ICIDS 2025 camera-ready version. This revision includes the final title, updated abstract, improved explanations of the narrative coherence framework, and minor editorial changes. Figures and examples have been refined for clarity. No new experiments were added

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12868 2025-11-18 cs.CV cs.AI 62%

Video Finetuning Improves Reasoning Between Frames

Ruiqi Yang, Tian Yun, Zihan Wang, Ellie Pavlick

机构 * Brown University(布朗大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at CogInterp @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11880 2025-11-18 cs.LG cs.AI cs.CV 62%

Transformers vs. Recurrent Models for Estimating Forest Gross Primary Production

David Montero, Miguel D. Mahecha, Francesco Martinuzzi, César Aybar, Anne Klosterhalfen, Alexander Knohl, Jesús Anaya, Clemens Mosig, Sebastian Wieneke

机构 * IEF, Leipzig University(莱比锡大学IEF) iDiv MPI PKS(马克斯·普朗克研究所) IPL, Universitat de València(瓦伦西亚大学IPL) Bioclimatology, University of Göttingen(哥廷根大学生物气候学系) GEMA, Universidad de Medellín(梅迪纳大学GEMA)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13190 2025-11-18 cs.CV 57%

Video Spatial Reasoning with Object-Centric 3D Rollout

Haoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang, Linglong Li, Ge Li, Xiaodan Liang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13035 2025-11-18 cs.LG cs.AI 57%

One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow

Zeyuan Wang, Da Li, Yulin Chen, Ye Shi, Liang Bai, Tianyuan Yu, Yanwei Fu

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted in AAAI 2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12291 2025-11-18 cs.CV 57%

One target to align them all: LiDAR, RGB and event cameras extrinsic calibration for Autonomous Driving

Andrea Bertogalli, Giacomo Boracchi, Luca Magri

机构 * DEIB Politecnico di Milano(都灵理工大学DEIB)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07251 2025-11-18 cs.CV 57%

Understanding Dynamic Scenes in Ego Centric 4D Point Clouds

Junsheng Huang, Shengyu Hao, Bocheng Hu, Hongwei Wang, Gaoang Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted as a poster to AAAI 2026; will be published in the proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02473 2025-11-18 cs.CV 57%

Generative Perception of Shape and Material from Differential Motion

Xinran Nicole Han, Ko Nishino, Todd Zickler

机构 * Harvard University(哈佛大学) Kyoto University(京都大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12154 2025-11-18 cs.LG cs.AI 57%

Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions

Gustavo Polleti, Marlesson Santana, Eduardo Fontes

机构 * Trustly

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12150 2025-11-18 cs.CV 57%

Breaking the Modality Wall: Time-step Mixup for Efficient Spiking Knowledge Transfer from Static to Event Domain

Yuqi Xie, Shuhan Ye, Yi Yu, Chong Wang, Qixin Zhang, Jiazhen Xu, Le Shen, Yuanbin Qian, Jiangbo Qian, Guoqi Li

机构 * Ningbo University(宁波大学) Nanyang Technological University(南洋理工大学) Merchants’ Guild Economics and Cultural(商帮经济与文化) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04953 2025-11-18 cs.CV 57%

APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval

Hong Gao, Yiming Bao, Xuezhen Tu, Bin Zhong, Linan Yue, Minling Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13419 2025-11-18 cs.LG physics.ao-ph 50%

MMWSTM-ADRAN+: A Novel Hybrid Deep Learning Architecture for Enhanced Climate Time Series Forecasting and Extreme Event Prediction

Shaheen Mohammed Saleh Ahmed, Hakan Hakan Guneyli

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09862 2025-11-18 cs.LG 50%

RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-Wave Point Cloud Sequence

Zengyuan Lai, Jiarui Yang, Songpengcheng Xia, Lizhou Lin, Lan Sun, Renwen Wang, Jianran Liu, Qi Wu, Ling Pei

专题命中 视频多模态 :cross-modal(abstract)

Comments Accepted by AAAI 2026 (extended version with supplementary materials)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 7 篇

2511.13545 2025-11-18 cs.CV cs.AI 84%

Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks

Md. Iqbal Hossain, Afia Sajeeda, Neeresh Kumar Perla, Ming Shao

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13189 2025-11-18 cs.CV cs.IR 79%

Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework

Diego Ortego, Marlon Rodríguez, Mario Almagro, Kunal Dahiya, David Jiménez, Juan C. SanMiguel

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments To appear at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08181 2025-11-18 cs.IR cs.AI 79%

MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System

Seung Hwan Cho, Yujin Yang, Danik Baeck, Minjoo Kim, Young-Min Kim, Heejung Lee, Sangjin Park

机构 * Department of Industrial Data Engineering, Hanyang University, Republic of Korea(工业数据工程系,翰阳大学) School of Interdisciplinary Industrial Studies, Hanyang University, Republic of Korea(跨学科工业研究学院,翰阳大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 2 figures, Accepted at RDGENAI at CIKM 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11050 2025-11-18 cs.CV cs.AI 73%

RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models

Yujin Wang, Quanfeng Liu, Jiaqi Fan, Jinlong Hong, Hongqing Chu, Mengjian Tian, Bingzhao Gao, Hong Chen

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10390 2025-11-18 cs.CV 70%

HIBMatch: Hypergraph Information Bottleneck for Semi-supervised Alzheimer's Progression

Zhongying Deng, Shujun Wang, Angelica I Aviles-Rivero, Zoe Kourtzi, Carola-Bibiane Schönlieb

机构 * Department of Applied Mathematics and Theoretical Physics, University of Cambridge(应用数学与理论物理系,剑桥大学) Department of Biomedical Engineering, The Hong Kong Polytechnic University(生物医学工程系,香港理工大学) Research Institute for Artificial Intelligence of Things, The Hong Kong Polytechnic University(物联网人工智能研究所,香港理工大学) Yau Mathematical Sciences Centre, Tsinghua University(叶德平数学科学中心,清华大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to the IEEE Journal of Biomedical and Health Informatics (To appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09754 2025-11-18 cs.LG cs.AI 57%

History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting

Sarthak Khanna, Armin Berger, Muskaan Chopra, David Berghaus, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫研究所-媒体工程部门) University of Bonn - Department of Computer Science(波恩大学-计算机科学系) Lamarr Institute for Machine Learning(拉马尔机器学习研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments Accepted in IEEE BigData 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11912 2025-11-18 cs.LG cs.CR 50%

A Systematic Study of Model Extraction Attacks on Graph Foundation Models

Haoyan Xu, Ruizhi Qian, Jiate Li, Yushun Dong, Minghao Lin, Hanson Yan, Zhengtao Yao, Qinghua Liu, Junhao Dong, Ruopeng Huang, Yue Zhao, Mengyuan Li

机构 * University of Southern California(南加州大学) Florida State University(佛罗里达州立大学) The Ohio State University(俄亥俄州立大学) Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 15 篇

2511.13647 2025-11-18 cs.CV 88%

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

Chunshi Wang, Junliang Ye, Yunhan Yang, Yang Li, Zizhuo Lin, Jun Zhu, Zhuo Chen, Yawei Luo, Chunchao Guo

机构 * Zhejiang University(浙江大学) Tencent Hunyuan(腾讯文言) Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13243 2025-11-18 cs.LG cs.AI cs.CV 84%

Uncovering and Mitigating Transient Blindness in Multimodal Model Editing

Xiaoqi Han, Ru Li, Ran Yi, Hongye Tan, Zhuomin Liang, Víctor Gutiérrez-Basulto, Jeff Z. Pan

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at AAAI'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13078 2025-11-18 cs.LG eess.AS eess.IV 79%

A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning

Liuyi Jin, Pasan Gunawardena, Amran Haroon, Runzhi Wang, Sangwoo Lee, Radu Stoleru, Michael Middleton, Zepeng Huo, Jeeeun Kim, Jason Moats

机构 * Computer Engineering, Texas A\&M University 3 Texas A\&M University Emergency Medical Services (EMS), 4 Biomedical Data Science, Stanford University, 5 Texas A\&M School of Public Health

专题命中 多模态生成 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏