arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4688 篇

2408.14842 2024-08-28 cs.CV cs.LG 88%

From Bias to Balance: Detecting Facial Expression Recognition Biases in Large Multimodal Foundation Models

Kaylee Chhua, Zhoujinyi Wen, Vedant Hathalia, Kevin Zhu, Sean O'Brien

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16510 2024-03-14 cs.CV 88%

Source-Free Domain Adaptation with Frozen Multimodal Foundation Model

Song Tang, Wenxin Su, Mao Ye, Xiatian Zhu

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08044 2024-01-22 cs.CV 88%

Benchmarking Robustness of Multimodal Image-Text Models under Distribution Shift

Jielin Qiu, Yi Zhu, Xingjian Shi, Florian Wenzel, Zhiqiang Tang, Ding Zhao, Bo Li, Mu Li

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV

Comments Accepted by Journal of Data-centric Machine Learning Research (DMLR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01064 2023-11-03 cs.CV cs.LG 88%

Multimodal Foundation Models for Zero-shot Animal Species Recognition in Camera Trap Images

Zalan Fabian, Zhongqi Miao, Chunyuan Li, Yuanhan Zhang, Ziwei Liu, Andrés Hernández, Andrés Montes-Rojas, Rafael Escucha, Laura Siabatto, Andrés Link, Pablo Arbeláez, Rahul Dodhia, Juan Lavista Ferres

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CV

Comments 18 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02903 2023-01-10 cs.LG cs.CV 88%

Transferring Pre-trained Multimodal Representations with Cross-modal Similarity Matching

Byoungjip Kim, Sungik Choi, Dasol Hwang, Moontae Lee, Honglak Lee

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

Comments 20 pages, 10 figures, NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07537 2022-10-17 cs.CV cs.LG 88%

Unconditional Image-Text Pair Generation with Multimodal Cross Quantizer

Hyungyung Lee, Sungjin Park, Joonseok Lee, Edward Choi

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CV

Comments BMVC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.06482 2022-09-21 cs.CL 88%

ITA: Image-Text Alignments for Multi-Modal Named Entity Recognition

Xinyu Wang, Min Gui, Yong Jiang, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu

专题命中 图文多模态 :multi-modal(title,abstract);image-text(title);cross-modal(abstract);分类 cs.CL

Comments Accepted to NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08009 2026-08-11 cs.CV cs.AI 新提交 88%

Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation

基于证据的取证推理:检测与定位多模态媒体篡改

Yichun Yeh, Yiheng Li, Xiaobo Hu, Zhen Lei, Yang Yang

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);image-text(abstract)

AI总结 本文针对多模态媒体篡改检测与定位问题,提出基于证据的取证推理框架,结合锚定-验证推理链、可验证奖励系统与模态解耦优势路由机制,实现最优性能与可解释性的统一。

Comments accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02597 2026-05-18 cs.CV cs.AI 88%

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs

视觉即理解:解锁因果注意力以实现模态互注意力用于多模态大语言模型

Wei-Yao Wang, Zhao Wang, Helen Suzuki, Yoshiyuki Kobayashi

机构 * Sony Group Corporation, Tokyo, Japan(索尼集团,日本东京)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI

AI总结 本文提出模态互注意力机制,通过解锁因果注意力,提升多模态理解性能,无需额外参数,在12个基准测试中平均提升6.2%。

Comments ICML 2026. Code is available at https://github.com/sony/aki

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06032 2026-07-08 cs.IR 新提交 88%

Uncertainty-Aware Cross-Modal Remote Sensing Image-Text Retrieval via Evidential Learning

通过证据学习实现不确定性感知的跨模态遥感图像-文本检索

Zhuoyue Wang, Xueqian Wang, Gang Li, Chengxi Li, Yongpan Liu, Yifang Ban

专题命中 图文多模态 :cross-modal(title,abstract);image-text(title,abstract)

AI总结 针对跨模态遥感图像-文本检索中测试条件非理想、现有方法检索结果不可靠的问题,提出基于证据学习的ELC方法,通过EDL、UCL、RL建模及优化,经实验验证该方法相比现有方法有竞争力且鲁棒性更强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17579 2023-05-04 cs.CL cs.AI cs.CV 88%

Multimodal Image-Text Matching Improves Retrieval-based Chest X-Ray Report Generation

Jaehwan Jeong, Katherine Tian, Andrew Li, Sina Hartung, Fardad Behzadi, Juan Calle, David Osayande, Michael Pohlen, Subathra Adithan, Pranav Rajpurkar

专题命中 图文多模态 :image-text(title,abstract);multimodal(title);分类 cs.CV、cs.CL、cs.AI

Journal ref Medical Imaging with Deep Learning 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26722 2026-08-28 cs.CV 新提交 87%

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

UniGeo:用于文本引导跨视角地理定位的多模态大语言模型

Jiahao Wen, Hang Yu, Zhedong Zheng

机构 * School of Computer Engineering and Science, Shanghai University(上海大学计算机工程与科学学院) Institute of Collaborative Innovation, University of Macau(澳门大学协同创新研究院)

专题命中 图文多模态 :multi-modal(title);MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract)

AI总结 UniGeo是一种统一多模态大语言模型,通过地理语义学习、跨视角生成及即插即用验证模块,在文本引导无人机地理定位任务中显著提升了检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03713 2026-08-12 cs.CV 版本更新 87%

Investigating Adversarial Robustness of Multi-modal Large Language Models

探究多模态大语言模型的对抗鲁棒性

Hashmat Shadab Malik, Muzammal Naseer, Salman Khan

机构 * Mohamed Bin Zayed University of AI, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋) Khalifa University, UAE(哈利法大学,阿联酋) Australian National University, Australia(澳大利亚国立大学,澳大利亚)

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 通过系统研究多模态大语言模型的对抗鲁棒性,提出诊断性CLIP对齐协议预测鲁棒视觉编码器的迁移效果,并证明端到端多模态对抗训练能显著提升模型在强对抗攻击下的性能。

Journal ref The 37th British Machine Vision Conference (BMVC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01215 2026-06-30 cs.CV cs.AI cs.CL cs.MM 87%

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs

将神经符号程序蒸馏到3D多模态大语言模型中

Wentao Mo, Yang Liu

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.CL、cs.AI

AI总结 提出APEIRIA,通过三阶段课程学习将符号推理模式蒸馏到3D多模态大语言模型中,实现透明推理与开放词汇空间推理的统一。

Comments To appear in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17243 2026-04-21 cs.CV 87%

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation

RemoteShield:为地球观测启用鲁棒的多模态大语言模型

Rui Min, Liang Yao, Shiyu Miao, Shengxiang Xu, Yuxuan Liu, Chuanyi Zhang, Shimin Di, Fan Liu

机构 * Hohai University(河海大学) Nanjing University(南京大学) Southeast University(东南大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);image-text(abstract);分类 cs.CV

AI总结 本文提出RemoteShield,一种针对地球观测的鲁棒多模态大语言模型,通过偏好学习提升在现实输入变化下的稳定性与一致性,实验显示其在多模态扰动下表现优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16822 2024-03-11 cs.CV 87%

EarthGPT: A Universal Multi-modal Large Language Model for Multi-sensor Image Comprehension in Remote Sensing Domain

Wei Zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, Xuerui Mao

专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract);cross-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11860 2023-11-28 cs.CV 87%

LION : Empowering Multimodal Large Language Model with Dual-Level Visual Knowledge

Gongwei Chen, Leyang Shen, Rui Shao, Xiang Deng, Liqiang Nie

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);image-text(abstract)

Comments Technical Report. Project page: https://rshaojimmy.github.io/Projects/JiuTian-LION Code: https://github.com/rshaojimmy/JiuTian

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10496 2022-04-29 cs.CV cs.AI cs.CL cs.LG cs.MM 87%

Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks

Zhecan Wang, Noel Codella, Yen-Chun Chen, Luowei Zhou, Xiyang Dai, Bin Xiao, Jianwei Yang, Haoxuan You, Kai-Wei Chang, Shih-fu Chang, Lu Yuan

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2201.05729

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15216 2026-07-17 cs.CV cs.AI 新提交 87%

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Symbal:检测模型生成字幕中的系统性对齐错误

Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari, Curtis Langlotz

专题命中 图文多模态 :MLLM(summary_cn,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 研究多模态大语言模型生成图像字幕时的系统性对齐错误检测问题,提出利用现成基础模型的结构化双阶段设置的Symbal方法,引入SymbalBench基准,Symbal在基准上表现出色,可辅助审核MLLM生成的字幕。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13805 2026-05-18 cs.CV cs.AI cs.LG 87%

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition

RAR:用于视觉识别的检索与排序增强多模态大语言模型

Ziyu Liu, Zeyi Sun, Yuhang Zang, Wei Li, Pan Zhang, Xiaoyi Dong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) MThreads, Inc.(MThreads公司) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :MLLM(summary_cn,abstract_cn);multimodal(abstract);multi-modal(abstract);image-text(abstract)

AI总结 RAR结合CLIP的检索与MLLM的排序,提升细粒度识别精度,增强少样本和零样本任务表现。

Comments Project: https://github.com/Liuziyu77/RAR

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16785 2026-04-21 cs.CV cs.AI 87%

Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games

弥合粗粒与细粒识别:一种混合方法用于交互教育游戏中的开放多粒度物体识别

Hanling Yi, Feng Lin, Mao Luo, Yifan Yang, Xiaotian Yu, Rong Xiao

机构 * Intellifusion Inc.(Intellifusion公司)

专题命中 图文多模态 :MLLM(summary_cn,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出HyMOR混合框架,结合MLLM和CLIP模型,解决开放多粒度物体识别中的粗粒与细粒任务差距问题,通过TBO数据集实验验证其在多模态内容生成和互动学习中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02595 2024-08-06 cs.CV cs.AI 87%

Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection

Sajal Aggarwal, Ananya Pandey, Dinesh Kumar Vishwakarma

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17174 2024-06-25 cs.CV cs.AI cs.LG 87%

Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution

Ying Wang, Tim G. J. Rudner, Andrew Gordon Wilson

专题命中 图文多模态 :multi-modal(title,abstract);image-text(title);分类 cs.CV、cs.AI

Comments Published in Advances in Neural Information Processing Systems 36 (NeurIPS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13288 2026-06-12 cs.CV cs.AI cs.CL 新提交 87%

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality

跨模态掩码组合概念建模以增强视觉-语言组合性

Wei Li, Zhen Huang, Xinmei Tian

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学,教育部脑启发智能感知与认知重点实验室) Independent Researcher(独立研究员)

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出MACCO框架,通过掩码一个模态的组合概念并从另一模态完整上下文重建,增强视觉-语言模型的组合理解能力,在五个基准上显著提升。

Comments Accepted to ACL 2026 Main Conference, 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29951 2026-05-29 cs.AI cs.CL cs.LG cs.MM 87%

MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

MuPHI: 通过语义基础奖励优化学习隐式多模态有害推理

Anisha Saha, Varsha Suresh, Teodora Kamova, Sophia Wiedmann, Timothy Hospedales, Vera Demberg

机构 * Max Planck Institute for Informatics(马克斯·普朗克院信息研究所) Saarland Informatics Campus(萨尔兰州信息校园) Saarland University(萨尔兰州大学) The University of Edinburgh(爱丁堡大学) Samsung AI Center, Cambridge(三星AI中心,剑桥)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 针对视觉语言模型在隐式跨模态有害语义推理上的不足,提出MuPHI数据集和MuPHIRM训练框架,通过多视角奖励优化联合语义学习,提升有害检测与推理质量及分布外鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12872 2026-05-14 cs.LG 87%

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning

SMA:用于数据高效多模态学习的子模ularity模态对齐器

Truong Pham, Anay Majee, Rishabh Iyer

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);multimodal foundation model(abstract)

AI总结 本文提出SMA,通过子模ularity目标提升多模态对齐效率,利用集合形式学习更丰富的跨模态结构,在低数据场景下实现显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20618 2026-01-29 cs.CV cs.AI cs.CL 87%

GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection

GDCNet:用于多模态讽刺检测的生成不一致比较网络

Shuguang Zhang, Junhong Lian, Guoxin Yu, Baoxun Xu, Xiang Ao

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(人工智能安全国家重点实验室,计算技术研究所(ICT),中国科学院(CAS)) University of Chinese Academy of Sciences, CAS(中国科学院大学) Pengcheng Laboratory(鹏城实验室) Shenzhen Stock Exchange(深圳证券交易所)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 GDCNet通过生成事实性图像描述和文本对比,提升多模态讽刺检测的准确性和鲁棒性。

Comments Accepted to 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18214 2025-12-04 cs.CV cs.AI cs.CL cs.LG 87%

VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety

VLSU:联合多模态理解在AI安全中的极限映射

Shruti Palaskar, Leon Gatys, Mona Abdelrahman, Mar Jacobo, Larry Lindsey, Rutika Moharir, Gunnar Lund, Yang Xu, Navid Shiee, Jeffrey Bigham, Charles Maalouf, Joseph Yitan Cheng

机构 * Apple(苹果公司)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 VLSU通过细粒度分类和组合分析,揭示了多模态安全评估中联合理解的缺陷,为改进AI安全研究提供关键测试平台。

Comments 10 pages, 5 figures, 4 tables, detailed appendix. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14714 2025-10-20 cs.CV cs.AI cs.CL 87%

KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection

Tuan-Vinh La, Minh-Hieu Nguyen, Minh-Son Dao

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Withdrawn by the authors due to lack of explicit agreement from all co-authors to post this version publicly on arXiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09623 2025-09-03 cs.CV cs.AI cs.CL 87%

Cross-Modal Adapter for Vision-Language Retrieval

Haojun Jiang, Jianke Zhang, Rui Huang, Chunjiang Ge, Zanlin Ni, Shiji Song, Gao Huang

机构 * Department of Automation, BNRist, Tsinghua University(自动化系,北京理工大学,清华大学) School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)

专题命中 图文多模态 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The accepted manuscript by Pattern Recognition 25 Journal. The published journal article is available at: https://doi.org/10.1016/j.patcog.2024.111144

Journal ref Pattern Recognition 2025

详情

展开后加载摘要…

URL PDF HTML 收藏