arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-27 至 2025-08-27 共收录 2 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 2 篇

2405.11793 2025-08-27 cs.CV 83%

MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise

Ruiqi Wu, Chenran Zhang, Jianle Zhang, Yi Zhou, Tao Zhou, Huazhu Fu

机构 * School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, China(教育部新一代人工智能技术及其交叉应用重点实验室) Nanjing University of Science and Technology, Nanjing, China(南京理工大学) Agency for Science, Technology and Research (A*STAR), Singapore(新加坡科技研究局)

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Early Accepted by The International Conference on Medical Image Computing and Computer Assisted Intervention(MICCAI)2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13152 2025-08-27 cs.CV cs.AI cs.RO 81%

SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models

Xiangyu Dong, Haoran Zhao, Jiang Gao, Haozhou Li, Xiaoguang Ma, Yaoming Zhou, Fuhai Chen, Juan Liu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏