arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-21 至 2025-08-21 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 3 篇

2508.08066 2025-08-21 cs.CV cs.AI cs.CL cs.LG 85%

ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model

Weitai Kang, Weiming Zhuang, Zhizhong Li, Yan Yan, Lingjuan Lyu

机构 * University of Illinois Chicago(伊利诺伊大学香槟分校) Sony AI(索尼人工智能)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 8 pages for the main paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14844 2025-08-21 cs.LG 78%

Multimodal Quantum Vision Transformer for Enzyme Commission Classification from Biochemical Representations

Murat Isik, Mandeep Kaur Saggi, Humaira Gowher, Sabre Kais

机构 * Purdue University(普渡大学) NC State University(北卡罗来纳州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at IEEE International Conference on Quantum Artificial Intelligence (QAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11762 2025-08-21 cs.LG cs.RO 78%

MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations

Daniel Bogdoll, Yitian Yang, Tim Joseph, Melih Yazgan, J. Marius Zöllner

机构 * FZI Research Center for Information Technology, Germany(德国弗赖堡信息科技研究中心) Karlsruhe Institute of Technology, Germany(德国卡尔斯鲁厄理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Daniel Bogdoll and Yitian Yang contributed equally. Accepted for publication at IV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏