arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-26 至 2025-09-26 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 9 篇

2508.06434 2025-09-26 cs.CV cs.AI 86%

CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment

Shengzhu Yang, Jiawei Du, Shuai Lu, Weihang Zhang, Ningli Wang, Huiqi Li

机构 * Beijing Institute of Technology(北京理工大学) Beijing Tongren Hospital(北京同仁医院)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07487 2025-09-26 cs.CV 85%

LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?

Bangyan Li, Wenxuan Huang, Zhenkun Gao, Yeqiang Wang, Yunhang Shen, Jingzhong Lin, Ling You, Yuxiang Shen, Shaohui Lin, Wanli Ouyang, Yuling Sun

机构 * East China Normal University(华东师范大学) The Chinese University of Hong Kong(香港中文大学) Northwest A&F University(西北农林科技大学) Tencent Youtu Lab(腾讯优图实验室) Xiamen University(厦门大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20961 2025-09-26 cs.CV cs.AI 84%

Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos

Sarmistha Das, R E Zera Marveen Lyngkhoi, Sriparna Saha, Alka Maurya

机构 * Indian Institute of Technology Patna(印度帕纳杰大学) CRISIL LTD(CRISIL公司)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20769 2025-09-26 cs.IR cs.AI cs.CV 81%

Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems

Tuo Zhang, Yuechun Sun, Ruiliang Liu

机构 * Museus University of Science and Technology of China(中国科学技术大学) British Museum(大英博物馆)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21251 2025-09-26 cs.CV cs.AI 76%

Instruction-tuned Self-Questioning Framework for Multimodal Reasoning

You-Won Jang, Yu-Jung Heo, Jaeseok Kim, Minsu Lee, Du-Seong Chang, Byoung-Tak Zhang

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments This paper was accepted to the "CLVL: 5th Workshop on Closing the Loop Between Vision and Language (ICCV 2023 CLVL workshop)."

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21287 2025-09-26 cs.CL cs.AI 73%

DisCoCLIP: A Distributional Compositional Tensor Network Encoder for Vision-Language Understanding

Kin Ian Lo, Hala Hawashin, Mina Abbaszadeh, Tilen Limback-Stokin, Hadi Wazni, Mehrnoosh Sadrzadeh

机构 * University College London(伦敦大学学院)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18174 2025-09-26 cs.CV cs.CL 73%

Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR

Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed, Ahmad Bastati, Zeina Aldallal, Sara Chrouf, Safwan AlModhayan

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00827 2025-09-26 cs.CV cs.AI 73%

IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves

Ruofan Wang, Juncheng Li, Yixu Wang, Bo Wang, Xiaosen Wang, Yan Teng, Yingchun Wang, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Huawei Technologies Ltd.(华为技术有限公司) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20792 2025-09-26 cs.CV cs.AI cs.LG 66%

DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation

Ved Umrajkar

机构 * Indian Institute of Technology, Roorkee(印度理工学院拉胡尔分校)

专题命中 图文多模态 :multimodal(abstract,comments);分类 cs.CV、cs.AI

Comments Accepted at ICCV2025 Workshop on Safe and Trustworthy Multimodal AI Systems

详情

展开后加载摘要…

URL PDF HTML 收藏