arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-04 至 2025-11-04 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 9 篇

2511.00716 2025-11-04 cs.LG 78%

Enhancing Heavy Rain Nowcasting with Multimodal Data: Integrating Radar and Satellite Observations

Rama Kassoumeh, David Rügamer, Henning Oppel

机构 * Bochum Institute of Technology(波恩技术学院) LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Okeanos Smart Data Solutions GmbH(Okeanos智能数据解决方案有限公司)

专题命中 视频多模态 :multimodal(title,abstract)

Comments accepted to ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13522 2025-11-04 cs.LG cs.NE q-fin.CP 78%

Cross-Modal Temporal Fusion for Financial Market Forecasting

Yunhua Pei, John Cartlidge, Anandadeep Mandal, Daniel Gold, Enrique Marcilio, Riccardo Mazzon

机构 * School of Computer Science, University of Bristol, Bristol, UK(布里斯托大学计算机科学学院) School of Engineering Mathematics and Technology, University of Bristol, Bristol, UK(布里斯托大学工程数学与技术学院) School of Engineering Mathematics(工程数学学院) Technology, University of Bristol, Bristol, UK(技术学院) Business School, University of Birmingham, Birmingham, UK(伯明翰大学商学院) Stratiphy Limited, London, UK(Stratiphy公司)

专题命中 视频多模态 :cross-modal(title,abstract)

Comments 10 pages, 4 figures, manuscript accepted to PAIS at ECAI-2025 European Conference on Artificial Intelligence, October 25-30, 2025, Bologna, Italy

Journal ref Frontiers in Artificial Intelligence and Applications, vol. 413, ECAI 2025, pp. 5360 - 5367

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01617 2025-11-04 cs.CV cs.IR 70%

Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers

Mohamed Eltahir, Ali Habibullah, Lama Ayash, Tanveer Hussain, Naeemullah Khan

机构 * King Abdullah University of Science and Technology (KAUST)(卡布斯大学) King Khalid University (KKU)(国王 Khalid 大学) Edge Hill University(埃德希尔大学)

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06350 2025-11-04 cs.CV 70%

Aligning Effective Tokens with Video Anomaly in Large Language Models

Yingxian Chen, Jiahui Liu, Ruidi Fan, Yanwei Li, Chirui Chang, Shizhen Zhao, Wilton W. T. Fok, Xiaojuan Qi, Yik-Chung Wu

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视频多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23603 2025-11-04 cs.CV 70%

PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity

Yuqian Yuan, Wenqiao Zhang, Xin Li, Shihao Wang, Kehan Li, Wentong Li, Jun Xiao, Lei Zhang, Beng Chin Ooi

机构 * Zhejiang University(浙江大学) DAMO Academy, Alibaba Group(阿里云图研究院) Hupan Lab(鸿篇实验室) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 22 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05783 2025-11-04 cs.CV cs.AI 62%

Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA

Zijie Song, Zhenzhen Hu, Yixiao Ma, Jia Li, Richang Hong

机构 * Hefei University of Technology, Hefei, China(合肥工业大学) University of Science and Technology of China, Hefei, China(中国科学技术大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01768 2025-11-04 cs.CV 57%

UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs

Zhe Liu, Jinghua Hou, Xiaoqing Ye, Jingdong Wang, Hengshuang Zhao, Xiang Bai

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology (HUST), Wuhan, China(华中科技大学电子信息与通信学院) School of Software Engineering, Huazhong University of Science and Technology (HUST), Wuhan, China(华中科技大学软件工程学院) Department of Computer Science, The University of Hong Kong, Hong Kong (HKU) SAR, China(香港大学计算机系) Baidu Inc., Beijing, China(百度公司)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00073 2025-11-04 cs.CV 57%

Habitat and Land Cover Change Detection in Alpine Protected Areas: A Comparison of AI Architectures

Harald Kristen, Daniel Kulmer, Manuela Hirschmugl

机构 * University of Graz(格拉茨大学) Joanneum Research(乔安姆研究)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15855 2025-11-04 q-bio.QM cs.AI cs.LG 57%

THFlow: A Temporally Hierarchical Flow Matching Framework for 3D Peptide Design

Dengdeng Huang, Shikui Tu

机构 * School of Computer Science Shanghai Jiao Tong University Shanghai, China(计算机科学学院 上海交通大学 上海中国)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏