arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4682 篇

2306.14182 2023-06-27 cs.CV cs.AI 84%

Switch-BERT: Learning to Model Multimodal Interactions by Switching Attention and Input

Qingpei Guo, Kaisheng Yao, Wei Chu

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by ECCV2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17415 2023-06-05 cs.CL cs.AI 84%

Exploring Better Text Image Translation with Multimodal Codebook

Zhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang, Jian Luan, Bin Wang, Degen Huang, Jinsong Su

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2023 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05432 2023-05-10 cs.CL cs.CV 84%

WikiWeb2M: A Page-Level Multimodal Wikipedia Dataset

Andrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown, Bryan A. Plummer, Kate Saenko, Jianmo Ni, Mandy Guo

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at the WikiWorkshop 2023. Data is readily available at https://github.com/google-research-datasets/wit/blob/main/wikiweb2m.md. arXiv admin note: text overlap with arXiv:2305.03668

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05313 2023-05-09 cs.CV cs.CL 84%

Refined Vision-Language Modeling for Fine-grained Multi-modal Pre-training

Lisai Zhang, Qingcai Chen, Zhijian Chen, Yunpeng Han, Zhonghua Li, Zhao Cao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Work in progress, v0.2

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13371 2023-05-03 cs.CV cs.MM 84%

Plug-and-Play Regulators for Image-Text Matching

Haiwen Diao, Ying Zhang, Wei Liu, Xiang Ruan, Huchuan Lu

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 13 pages, 9 figures, Accepted by TIP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13455 2023-03-24 cs.CV cs.CL 84%

CoBIT: A Contrastive Bi-directional Image-Text Generation Model

Haoxuan You, Mandy Guo, Zhecan Wang, Kai-Wei Chang, Jason Baldridge, Jiahui Yu

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11090 2023-03-21 cs.CV cs.AI 84%

Scene Graph Based Fusion Network For Image-Text Retrieval

Guoliang Wang, Yanlei Shang, Yong Chen

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08281 2022-12-19 cs.CV cs.MM 84%

HGAN: Hierarchical Graph Alignment Network for Image-Text Retrieval

Jie Guo, Meiting Wang, Yan Zhou, Bin Song, Yuhao Chi, Wei Fan, Jianglong Chang

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03853 2022-08-30 cs.IR cs.CL cs.CV 84%

Where Does the Performance Improvement Come From? -- A Reproducibility Concern about Image-Text Retrieval

Jun Rao, Fei Wang, Liang Ding, Shuhan Qi, Yibing Zhan, Weifeng Liu, Dacheng Tao

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments SIGIR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01917 2022-06-15 cs.CV cs.LG cs.MM 84%

CoCa: Contrastive Captioners are Image-Text Foundation Models

Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, Yonghui Wu

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12105 2022-06-01 cs.CV cs.CL 84%

HiVLP: Hierarchical Vision-Language Pre-Training for Fast Image-Text Retrieval

Feilong Chen, Xiuyi Chen, Jiaxin Shi, Duzhen Zhang, Jianlong Chang, Qi Tian

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08594 2022-05-04 cs.CV cs.CL 84%

Twitter-COMMs: Detecting Climate, COVID, and Military Multimodal Misinformation

Giscard Biamby, Grace Luo, Trevor Darrell, Anna Rohrbach

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05587 2021-12-16 cs.CV cs.CL cs.LG 84%

Unified Multimodal Pre-training and Prompt-based Tuning for Vision-Language Understanding and Generation

Tianyi Liu, Zuxuan Wu, Wenhan Xiong, Jingjing Chen, Yu-Gang Jiang

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02114 2021-11-04 cs.CV cs.CL cs.LG 84%

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, Aran Komatsuzaki

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Short version. Accepted at Data Centric AI NeurIPS Workshop 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04699 2021-09-23 cs.CL cs.CV cs.LG 84%

EfficientCLIP: Efficient Cross-Modal Pre-training by Ensemble Confident Learning and Language Modeling

Jue Wang, Haofan Wang, Jincan Deng, Weijia Wu, Debing Zhang

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.11735 2020-11-25 cs.AI cs.CV cs.LG 84%

Large Scale Multimodal Classification Using an Ensemble of Transformer Models and Co-Attention

Varnith Chordia, Vijay Kumar BG

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.11550 2020-10-23 cs.CV cs.MM 84%

Learning Dual Semantic Relations with Graph Attention for Image-Text Matching

Keyu Wen, Xiaodong Gu, Qingrong Cheng

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 14pages, 9 figures. Accepted at: IEEE Transactions on Circuits and Systems for Video Technology (Early Access Print) | |Codes Available at: https://github.com/kywen1119/DSRAN

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.12070 2020-10-13 cs.CV cs.CL cs.LG 84%

Deep Multimodal Neural Architecture Search

Zhou Yu, Yuhao Cui, Jun Yu, Meng Wang, Dacheng Tao, Qi Tian

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accept to ACM MM2020, code available at https://github.com/MILVLG/mmnas/

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.01473 2020-03-05 cs.CL cs.CV cs.LG 84%

XGPT: Cross-modal Generative Pre-Training for Image Captioning

Qiaolin Xia, Haoyang Huang, Nan Duan, Dongdong Zhang, Lei Ji, Zhifang Sui, Edward Cui, Taroon Bharti, Xin Liu, Ming Zhou

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 12 pages, 3 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1411.2539 2014-11-11 cs.LG cs.CL cs.CV 84%

Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Ryan Kiros, Ruslan Salakhutdinov, Richard S. Zemel

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 13 pages. NIPS 2014 deep learning workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15243 2025-09-22 cs.CV 84%

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

Muhammad Imran, Yugyung Lee

机构 * Computer Science, School of Science and Engineering, University of Missouri - Kansas City(计算机科学系,科学与工程学院,密苏里大学-堪萨斯城分校)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV;multimodal(journal_ref)

Comments 8 pages, 6 figures, 3 tables

Journal ref Non-Archival track - The First Workshop on Multimodal Knowledge and Language Modeling IJCAI 2025 Workshop, August 16, 2025 IJCAI 2025 Workshop, August 16, 2025 Room 516B, Palais des congrès, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12595 2024-10-17 cs.CV 84%

CMAL: A Novel Cross-Modal Associative Learning Framework for Vision-Language Pre-Training

Zhiyuan Ma, Jianjun Li, Guohui Li, Kaiyan Huang

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments vision-language pre-training, contrastive learning, cross-modal, associative learning, associative mapping classification

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23570 2026-08-26 cs.CL 新提交 83%

Taming Visual Neglect: A Variational Information Bottleneck Framework for Adaptive Attention in Multimodal In-Context Learning

驯服视觉忽视:用于多模态上下文学习中自适应注意力的变分信息瓶颈框架

Kaito Tanaka, Yuji Nishimura, Keisuke Matsuda, Aya Nakayama

机构 * SANNO University(山王大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 针对多模态上下文学习中视觉上下文时而被利用时而被忽视的问题,提出VIB-ICL框架,通过CMIG量化跨模态信息,推导理论界并经五组基准实验验证,实现准确率提升与演示样本减少。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21133 2026-08-24 cs.CV cs.CR 新提交 83%

Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI

仅掩码不足:面向医疗AI的多模态去标识化生成式修复

Shiva Shrestha, Zongxing Xie, Chen Zhao, Liran Ma, Zhipeng Cai, Honghui Xu

机构 * Kennesaw State University(肯尼索州立大学) Miami University(迈阿密大学) Georgia State University(佐治亚州立大学)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 针对医疗图像-文本数据的PHI泄漏问题,提出端到端多模态净化框架ClinX,结合图像侧生成式修复与文本侧渐进式去标识化,在MedVQA中验证其优于仅OCR掩码的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19207 2026-08-21 cs.CL 新提交 83%

Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages

合规性、能力与冲突:基于系统消息的多模态大语言模型基准测试

Juan Yeo, Geewook Kim

机构 * NAVER Cloud(NAVER云) KAIST AI(韩国科学技术院人工智能)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 该研究构建基准VSysBench测试多模态大语言模型在系统消息下的合规性,发现施加系统消息会降低任务准确率,开放权重模型易受用户冲突影响,视觉约束最难。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17450 2026-08-13 cs.IR cs.AI 版本更新 83%

VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation

VLM2Rec: 解决视觉语言模型嵌入器在多模态序列推荐中的模态崩溃问题

Junyoung Kim, Woojoo Kim, Wonbin Kweon, Jaehyung Lim, Dongha Kim, Hwanjo Yu

机构 * Pohang University of Science

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出VLM2Rec框架,通过弱模态惩罚对比学习和跨模态关系拓扑正则化,解决多模态序列推荐中模态崩溃问题,提升推荐准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09772 2026-08-11 cs.CL 新提交 83%

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

PragMatch:分离大视觉语言模型中的语用不一致与跨模态不匹配

Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CL

AI总结 本研究提出PragMatch基准,揭示大视觉语言模型易受表面线索影响,为评估多模态语用推理提供测试平台。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08060 2026-08-11 cs.CV 新提交 83%

ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models

ZOMP:面向视觉-语言模型的零阶多模态提示调优

Sajjad Ghiasvand, Yifan Yang, Mahnoosh Alizadeh, Ramtin Pedarsani

机构 * UC Santa Barbara(加州大学圣巴巴拉分校)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出ZOMP方法,通过跨模态低秩重参数化等技术,在仅前向传递的条件下高效调优CLIP模型,在13个基准上优于现有无反向传播的提示调优方法,泛化能力更强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06939 2026-08-11 cs.CV 版本更新 83%

Degradation-Aware Prompt Learning with Cross-Modal Compensation for Adverse Weather Removal

面向恶劣天气去除的退化感知跨模态补偿提示学习

Wanshu Fan, Yunzhe Zhang, Yue Shen, Liyan Wang, Jing Qin, Kin-Man Lam, Cong Wang, Jinshan Pan

机构 * School of Software Engineering, Dalian University(大连大学软件工程学院) School of Mathematical Sciences, Dalian University of Technology(大连理工大学数学科学学院) The Hong Kong Polytechnic University(香港理工大学) University of California, San Francisco(加利福尼亚大学旧金山分校) School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 该研究针对恶劣天气导致图像退化影响视觉系统可靠性的问题,提出 DCMPC-Net 模型,通过跨模态提示补偿实现鲁棒的恶劣天气图像修复,性能优于现有最先进方法。

Comments Accepted for publication in IEEE Transactions on Image Processing. The code is available at: https://github.com/fanamber831/DCMPC-Net

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02048 2026-08-11 cs.CV 版本更新 83%

Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models

Jagle:构建大规模日语多模态预训练数据集以构建视觉-语言模型

Issa Sugiura, Keito Sasagawa, Keisuke Nakao, Koki Maeda, Ziqi Yin, Zhishen Yang, Shuhei Kurita, Yusuke Oda, Ryoko Tokuhisa, Daisuke Kawahara, Naoaki Okazaki

机构 * Kyoto University(京都大学) NII LLMC(国立信息学研究所LLMC) Waseda University(早稻田大学) Institute of Science Tokyo(东京科学大学) NII(国立信息学研究所) Aichi Institute of Technology(爱知工业大学) Institute of Physical and Chemical Research(理化学研究所)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出Jagle,一个包含920万实例的日语多模态预训练数据集,通过多种策略生成VQA对,实验表明其在日语任务中表现优异,且与FineVision结合能提升英文性能。

Comments Accepted to COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏