arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-02 至 2025-10-02 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 6 篇

2506.23972 2025-10-02 cs.CV 83%

Learning Frequency and Memory-Aware Prompts for Multi-Modal Object Tracking

Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming zhou, Gangshan Wu, Jinde Cao

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) School of Information Science and Engineering(信息科学与工程学院) School of Mathematics(数学学院) Purple Mountain Laboratories(紫金山实验室)

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22646 2025-10-02 cs.CV cs.AI cs.CL 82%

Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs

Xingyu Fu, Siyi Liu, Yinuo Xu, Pan Lu, Guangqiuse Hu, Tianbo Yang, Taran Anantasagar, Christopher Shen, Yikai Mao, Yuanzhe Liu, Keyush Shah, Chung Un Lee, Yejin Choi, James Zou, Dan Roth, Chris Callison-Burch

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://deeptracereward.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25393 2025-10-02 cs.CV cs.AI 81%

Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction

Wendong Yao, Binhua Huang, Soumyabrata Dev

机构 * ADAPT SFI Research Centre, School of Computer Science, University College Dublin(ADAPT SFI研究所以及计算机科学学院,都柏林大学学院)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments This paper is submitted to IEEE Transactions on Geoscience and Remote Sensing for reviewing

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17481 2025-10-02 eess.SP cs.AI cs.LG 79%

Toward Foundational Model for Sleep Analysis Using a Multimodal Hybrid Self-Supervised Learning Framework

Cheol-Hui Lee, Hakseung Kim, Byung C. Yoon, Dong-Joo Kim

机构 * Department of Brain and Cognitive Engineering, Korea University(脑科学与认知工程系,韩国大学) Interdisciplinary Program in Precision Public Health, Korea University(精准公共卫生跨学科项目,韩国大学) Department of Radiology, Stanford University School of Medicine(放射科,斯坦福大学医学院) VA Palo Alto Health Care System(帕洛阿尔托医疗系统) Department of Neurology, Korea University College of Medicine(神经病学系,韩国大学医学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments 18 pages, 5 figures

Journal ref IEEE Transactions on Cybernetics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19778 2025-10-02 cs.AI 79%

Multimodal Large Language Models for Bioimage Analysis

Shanghang Zhang, Gaole Dai, Tiejun Huang, Jianxu Chen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04528 2025-10-02 cs.LG 50%

Federated Dynamic Modeling and Learning for Spatiotemporal Data Forecasting

Thien Pham, Angelo Furno, Faïcel Chamroukhi, Latifa Oukhellou

机构 * COSYS-GRETTIA, Gustave Eiffel University, 77420 France(COSYS-GRETTIA,巴黎-伊夫林大学) ENTPE, University of Lyon(ENTPE,里昂大学) the LICIT-ECO7 University Gustave Eiffel, France(LICIT-ECO7 巴黎-伊夫林大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏