arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-12 至 2025-08-12 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 9 篇

2508.06939 2025-08-12 cs.AI cs.LG 79%

Intrinsic Explainability of Multimodal Learning for Crop Yield Prediction

Hiba Najjar, Deepak Pathak, Marlon Nuske, Andreas Dengel

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07401 2025-08-12 cs.CV 70%

LET-US: Long Event-Text Understanding of Scenes

Rui Chen, Xingyu Chen, Shaoan Wang, Shihan Kong, Junzhi Yu

专题命中 视频多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08035 2025-08-12 cs.CV cs.AI 62%

LVBench: An Extreme Long Video Understanding Benchmark

Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Xiaotao Gu, Shiyu Huang, Bin Xu, Yuxiao Dong, Ming Ding, Jie Tang

机构 * Zhipu AI(智谱AI) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07989 2025-08-12 cs.CV cs.HC 57%

The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility

Xiantao Zhang

机构 * Beihang University(北航大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 9 pages, 3 figures, 2 tables. Accepted at CV4A11y, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07626 2025-08-12 cs.CV cs.RO 57%

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning

Dejie Yang, Zijing Zhao, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学计算机技术研究院) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07312 2025-08-12 cs.CV 57%

MobileViCLIP: An Efficient Video-Text Model for Mobile Devices

Min Yang, Zihan Jia, Zhilin Dai, Sheng Guo, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) MyBank, Ant Group(蚂蚁集团MyBank) Shanghai AI Lab(上海AI实验室)

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07006 2025-08-12 eess.IV cs.CV 57%

Spatio-Temporal Conditional Diffusion Models for Forecasting Future Multiple Sclerosis Lesion Masks Conditioned on Treatments

Gian Mario Favero, Ge Ya Luo, Nima Fathi, Justin Szeto, Douglas L. Arnold, Brennan Nichyporuk, Chris Pal, Tal Arbel

机构 * McGill University(麦吉尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to MICCAI 2025 (LMID Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12524 2025-08-12 cs.CV cs.HC cs.LG eess.IV 57%

Inference-Time Gaze Refinement for Micro-Expression Recognition: Enhancing Event-Based Eye Tracking with Motion-Aware Post-Processing

Nuwan Bandara, Thivya Kandappu, Archan Misra

机构 * School of Computing(计算学院) Information Systems, Singapore Management University, Singapore(信息系统,新加坡管理大学,新加坡)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at 4DMR@IJCAI25: International IJCAI Workshop on 1st Challenge and Workshop for 4D Micro-Expression Recognition for Mind Reading, August 29, 2025, Guangzhou, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06892 2025-08-12 astro-ph.SR physics.space-ph 50%

Large Model Driven Solar Activity AI Forecaster: A Scalable Dual Data-Model Framework

Jingjing Wang, Pengyu Liang, Tingyu Wang, Ming Li, Yanmei Cui, Siwei Liu, Xin Huang, Xiang Li, Minghui Zhang, Yunshi Zeng, Zhu Cao, Jiekang Feng, Qinghua Hu, Bingxian Luo, Bing Cao

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏