arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-18 至 2025-11-18 共收录 22 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 22 篇

2511.12460 2025-11-18 cs.LG cs.AI 83%

Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network for Multimodal Depression Detection

Changzeng Fu, Shiwen Zhao, Yunze Zhang, Zhongquan Jian, Shiqi Zhao, Chaoran Liu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments AAAI 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04351 2025-11-18 eess.SP 82%

RCMCL: A Unified Contrastive Learning Framework for Robust Multi-Modal (RGB-D, Skeleton, Point Cloud) Action Understanding

Hasan Akgul, Mari Eplik, Javier Rojas, Akira Yamamoto, Rajesh Kumar, Maya Singh

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)

Comments 11 pages, 6 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13655 2025-11-18 cs.CV cs.LG 79%

OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation

Henry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng, Joseph Redmon, Hadrien Sablon, Ryan Park, Jacob Morrison, Alexandra Buraczynski, Karen Farley, Joshua Hansen, Andrew Howe, Patrick Alan Johnson, Mark Otterlee, Ted Schmitt, Hunter Pitelka, Stephen Daspit, Rachel Ratner, Christopher Wilhelm, Sebastian Wood, Mike Jacobi, Hannah Kerner, Evan Shelhamer, Ali Farhadi, Ranjay Krishna, Patrick Beukema

机构 * Allen Institute for AI(人工智能研究所) University of Washington(华盛顿大学) Arizona State University(亚利桑那州立大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12196 2025-11-18 cs.CV cs.HC 79%

Cross-View Cross-Modal Unsupervised Domain Adaptation for Driver Monitoring System

Aditi Bhalla, Christian Hellert, Enkelejda Kasneci

机构 * School of Social Sciences and Technology(社会科学与技术学院)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12027 2025-11-18 cs.CV cs.AI 73%

GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory

Jeong Hun Yeo, Sangyun Chung, Sungjune Park, Dae Hoe Kim, Jinyoung Moon, Yong Man Ro

机构 * Integrated Vision and Language Lab., School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(整合视觉与语言实验室,电气工程学院,韩国科学技术院(KAIST)) Visual Intelligence Research Section, Superintelligence Creative Research Laboratory, Electronics and Telecommunications Research Institute (ETRI)(视觉智能研究部,超智能创意研究实验室,电子电信研究院)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11908 2025-11-18 cs.CV cs.AI 73%

PI-NAIM: Path-Integrated Neural Adaptive Imputation Model

Afifa Khaled, Ebrahim Hamid Sumiea

机构 * University of Science and Technology of China(中国科学技术大学) Universiti Teknologi PETRONAS(Petronas科技大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13054 2025-11-18 cs.CV 70%

ViSS-R1: Self-Supervised Reinforcement Video Reasoning

Bo Fang, Yuxin Song, Qiangqiang Wu, Haoyuan Sun, Wenhao Wu, Antoni B. Chan

机构 * City University of Hong Kong(香港城市大学) Baidu Inc.(百度公司) Tsinghua University(清华大学) The University of Sydney(悉尼大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Our paper was initially titled "Video-SSR1: Self-Supervised Reinforcement Video Reasoning." Upon noticing its close resemblance to the title of a recently released paper, we have decided to rename our work as "ViSS-R1."

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11576 2025-11-18 cs.CV 70%

Causality Matters: How Temporal Information Emerges in Video Language Models

Yumeng Shi, Quanyu Long, Yin Wu, Wenya Wang

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19002 2025-11-18 cs.CV cs.AI cs.CL cs.LG 67%

VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

Hao Wang, Eiki Murata, Lingfang Zhang, Ayako Sato, So Fukuda, Ziqi Yin, Wentao Hu, Keisuke Nakao, Yusuke Nakamura, Sebastian Zwirner, Yi-Chia Chen, Hiroyuki Otomo, Hiroki Ouchi, Daisuke Kawahara

机构 * Waseda University(早稻田大学) CyberAgent, Inc.(CyberAgent公司) AI Shift, Inc.(AI Shift公司) Nara Institute of Science and Technology(奈良研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10008 2025-11-18 cs.MM cs.AI cs.CV 67%

Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives

Yi-Chun Chen

机构 * Yale University(耶鲁大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Updated with the ICIDS 2025 camera-ready version. This revision includes the final title, updated abstract, improved explanations of the narrative coherence framework, and minor editorial changes. Figures and examples have been refined for clarity. No new experiments were added

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12868 2025-11-18 cs.CV cs.AI 62%

Video Finetuning Improves Reasoning Between Frames

Ruiqi Yang, Tian Yun, Zihan Wang, Ellie Pavlick

机构 * Brown University(布朗大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at CogInterp @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11880 2025-11-18 cs.LG cs.AI cs.CV 62%

Transformers vs. Recurrent Models for Estimating Forest Gross Primary Production

David Montero, Miguel D. Mahecha, Francesco Martinuzzi, César Aybar, Anne Klosterhalfen, Alexander Knohl, Jesús Anaya, Clemens Mosig, Sebastian Wieneke

机构 * IEF, Leipzig University(莱比锡大学IEF) iDiv MPI PKS(马克斯·普朗克研究所) IPL, Universitat de València(瓦伦西亚大学IPL) Bioclimatology, University of Göttingen(哥廷根大学生物气候学系) GEMA, Universidad de Medellín(梅迪纳大学GEMA)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13190 2025-11-18 cs.CV 57%

Video Spatial Reasoning with Object-Centric 3D Rollout

Haoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang, Linglong Li, Ge Li, Xiaodan Liang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13035 2025-11-18 cs.LG cs.AI 57%

One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow

Zeyuan Wang, Da Li, Yulin Chen, Ye Shi, Liang Bai, Tianyuan Yu, Yanwei Fu

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted in AAAI 2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12291 2025-11-18 cs.CV 57%

One target to align them all: LiDAR, RGB and event cameras extrinsic calibration for Autonomous Driving

Andrea Bertogalli, Giacomo Boracchi, Luca Magri

机构 * DEIB Politecnico di Milano(都灵理工大学DEIB)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07251 2025-11-18 cs.CV 57%

Understanding Dynamic Scenes in Ego Centric 4D Point Clouds

Junsheng Huang, Shengyu Hao, Bocheng Hu, Hongwei Wang, Gaoang Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted as a poster to AAAI 2026; will be published in the proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02473 2025-11-18 cs.CV 57%

Generative Perception of Shape and Material from Differential Motion

Xinran Nicole Han, Ko Nishino, Todd Zickler

机构 * Harvard University(哈佛大学) Kyoto University(京都大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12154 2025-11-18 cs.LG cs.AI 57%

Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions

Gustavo Polleti, Marlesson Santana, Eduardo Fontes

机构 * Trustly

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12150 2025-11-18 cs.CV 57%

Breaking the Modality Wall: Time-step Mixup for Efficient Spiking Knowledge Transfer from Static to Event Domain

Yuqi Xie, Shuhan Ye, Yi Yu, Chong Wang, Qixin Zhang, Jiazhen Xu, Le Shen, Yuanbin Qian, Jiangbo Qian, Guoqi Li

机构 * Ningbo University(宁波大学) Nanyang Technological University(南洋理工大学) Merchants’ Guild Economics and Cultural(商帮经济与文化) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04953 2025-11-18 cs.CV 57%

APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval

Hong Gao, Yiming Bao, Xuezhen Tu, Bin Zhong, Linan Yue, Minling Zhang

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13419 2025-11-18 cs.LG physics.ao-ph 50%

MMWSTM-ADRAN+: A Novel Hybrid Deep Learning Architecture for Enhanced Climate Time Series Forecasting and Extreme Event Prediction

Shaheen Mohammed Saleh Ahmed, Hakan Hakan Guneyli

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09862 2025-11-18 cs.LG 50%

RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-Wave Point Cloud Sequence

Zengyuan Lai, Jiarui Yang, Songpengcheng Xia, Lizhou Lin, Lan Sun, Renwen Wang, Jianran Liu, Qi Wu, Ling Pei

专题命中 视频多模态 :cross-modal(abstract)

Comments Accepted by AAAI 2026 (extended version with supplementary materials)

详情

展开后加载摘要…

URL PDF HTML 收藏