arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-21 至 2025-08-21 共收录 39 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 1 篇

2508.14055 2025-08-21 cs.CL cs.AI 62%

T-REX: Table -- Refute or Entail eXplainer

Tim Luka Horstmann, Baptiste Geisenberger, Mehwish Alam

机构 * Institut Polytechnique de Paris(巴黎政治科技学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 3 篇

2508.08066 2025-08-21 cs.CV cs.AI cs.CL cs.LG 85%

ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model

Weitai Kang, Weiming Zhuang, Zhizhong Li, Yan Yan, Lingjuan Lyu

机构 * University of Illinois Chicago(伊利诺伊大学香槟分校) Sony AI(索尼人工智能)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 8 pages for the main paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14844 2025-08-21 cs.LG 78%

Multimodal Quantum Vision Transformer for Enzyme Commission Classification from Biochemical Representations

Murat Isik, Mandeep Kaur Saggi, Humaira Gowher, Sabre Kais

机构 * Purdue University(普渡大学) NC State University(北卡罗来纳州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at IEEE International Conference on Quantum Artificial Intelligence (QAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11762 2025-08-21 cs.LG cs.RO 78%

MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations

Daniel Bogdoll, Yitian Yang, Tim Joseph, Melih Yazgan, J. Marius Zöllner

机构 * FZI Research Center for Information Technology, Germany(德国弗赖堡信息科技研究中心) Karlsruhe Institute of Technology, Germany(德国卡尔斯鲁厄理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Daniel Bogdoll and Yitian Yang contributed equally. Accepted for publication at IV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 5 篇

2508.14504 2025-08-21 cs.CV cs.AI 87%

PB-IAD: Utilizing multimodal foundation models for semantic industrial anomaly detection in dynamic manufacturing environments

Bernd Hofmann, Albert Scheck, Joerg Franke, Patrick Bruendl

专题命中 其他多模态 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14523 2025-08-21 cs.LG 78%

Great GATsBi: Hybrid, Multimodal, Trajectory Forecasting for Bicycles using Anticipation Mechanism

Kevin Riehl, Shaimaa K. El-Baklish, Anastasios Kouvelas, Michail A. Makridis

机构 * Traffic Engineering Group, Institute for Transport Planning and Systems, ETH Zürich(交通规划与系统研究所,瑞士苏黎世联邦理工学院)

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14363 2025-08-21 physics.med-ph 78%

Multimodal concentric surface coils for enhanced sensitivity in MR imaging

Yunkun Zhao, Aditya A Bhosale, Xiaoliang Zhang

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17976 2025-08-21 cs.SE cs.AI 57%

The importance of visual modelling languages in generative software engineering

Roberto Rossi

机构 * Business School, University of Edinburgh(爱丁堡大学商学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 9 pages, working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14190 2025-08-21 cs.CR cs.CL cs.LG 57%

Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text

Zixin Rao, Youssef Mohamed, Shang Liu, Zeyan Liu

机构 * University of Georgia(佐治亚大学) Egypt-Japan University of Science and Technology(埃及-日本科学技术大学) University of Louisville(路易斯维尔大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CL

Comments Securecomm 2025

详情

展开后加载摘要…

URL PDF HTML 收藏