arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46294 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4676 篇

2002.12585 2020-03-17 cs.CV cs.CL cs.LG 81%

Exploring and Distilling Cross-Modal Information for Image Captioning

Fenglin Liu, Xuancheng Ren, Yuanxin Liu, Kai Lei, Xu Sun

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by IJCAI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.08622 2020-01-09 cs.CV cs.CL cs.LG stat.ML 81%

Variational Hetero-Encoder Randomized GANs for Joint Image-Text Modeling

Hao Zhang, Bo Chen, Long Tian, Zhengjue Wang, Mingyuan Zhou

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.CL

Comments ICLR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.00212 2019-11-04 cs.LG cs.CL cs.CV stat.ML 81%

Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning

Tao Jin, Siyu Huang, Yingming Li, Zhongfei Zhang

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted as a long paper at EMNLP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.09953 2019-09-27 cs.CV cs.AI 81%

Learning Visual Relation Priors for Image-Text Matching and Image Captioning with Neural Scene Graph Generators

Kuang-Huei Lee, Hamid Palangi, Xi Chen, Houdong Hu, Jianfeng Gao

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.13358 2019-06-03 cs.CL cs.CV 81%

Multi-modal Discriminative Model for Vision-and-Language Navigation

Haoshuo Huang, Vihan Jain, Harsh Mehta, Jason Baldridge, Eugene Ie

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at SpLU-RoboNLP 2019 (workshop at NAACL)

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.12980 2019-05-31 cs.CV cs.CL 81%

Interactive-predictive neural multimodal systems

Álvaro Peris, Francisco Casacuberta

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments To appear at IbPRIA 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.08024 2018-07-24 cs.CV cs.AI cs.LG 81%

Stacked Cross Attention for Image-Text Matching

Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, Xiaodong He

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ECCV 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06786 2018-05-25 cs.CL cs.CV cs.IR 81%

Quantifying the visual concreteness of words and topics in multimodal datasets

Jack Hessel, David Mimno, Lillian Lee

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments NAACL HLT 2018, 14 pages, 6 figures, data available at http://www.cs.cornell.edu/~jhessel/concreteness/concreteness.html

Journal ref 2018 North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT)

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.05038 2017-09-18 cs.CV cs.CL cs.LG 81%

Self-Guiding Multimodal LSTM - when we do not have a perfect training dataset for image captioning

Yang Xian, Yingli Tian

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments The paper is under consideration at Computer Vision and Image Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.07601 2017-07-25 cs.CL cs.CV 81%

Image Pivoting for Learning Multilingual Multimodal Representations

Spandana Gella, Rico Sennrich, Frank Keller, Mirella Lapata

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 7 pages, EMNLP 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1412.6632 2015-06-12 cs.CV cs.CL cs.LG 81%

Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, Alan Yuille

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Add a simple strategy to boost the performance of image captioning task significantly. More details are shown in Section 8 of the paper. The code and related data are available at https://github.com/mjhucla/mRNN-CR ;. arXiv admin note: substantial text overlap with arXiv:1410.1090

Journal ref ICLR 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01664 2026-08-04 cs.CV cs.LG 新提交 80%

FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering

FAU参加2026年ImageCLEF多模态推理任务:鲁棒候选评分与简洁多语种视觉问答

Mohamed Basem, Vincent Christlein

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗里德里希-亚历山大-埃尔兰根-纽伦堡大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 FAU提交的多模态推理系统,未做任务特定模型训练,在2026年ImageCLEF竞赛中,获Visual MCQ第三名、Visual OpenQA第一名,凸显推理工程的实用价值。

Comments 16 pages, 3 figures, 7 tables. CLEF 2026 Working Notes, ImageCLEF 2026 Multimodal Reasoning Task

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17582 2025-11-13 cs.HC cs.GR cs.MM 80%

Crafting Dynamic Virtual Activities with Advanced Multimodal Models

Changyang Li, Qingan Yan, Minyoung Kim, Zhan Li, Yi Xu, Lap-Fai Yu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.MM

Journal ref C. Li, Q. Yan, M. Kim, Z. Li, Y. Xu and L. -F. Yu, "Crafting Dynamic Virtual Activities with Advanced Multimodal Models," 2025 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 120-130

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19196 2025-08-19 cs.RO cs.CL cs.HC 80%

Towards Multimodal Social Conversations with Robots: Using Vision-Language Models

Ruben Janssens, Tony Belpaeme

机构 * Ghent University–imec(根特大学–imec)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at the workshop "Human - Foundation Models Interaction: A Focus On Multimodal Information" (FoMo-HRI) at IEEE RO-MAN 2025 (Camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09914 2025-04-15 cs.CV 80%

Improving Multimodal Hateful Meme Detection Exploiting LMM-Generated Knowledge

Maria Tzelepi, Vasileios Mezaris

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted for publication, Multimodal Learning and Applications Workshop (MULA 2025) @ IEEE/CVF CVPR 2025, Nashville, TN, USA, June 2025. This is the authors' "accepted version"

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03452 2024-08-05 cs.CV cs.LG 80%

Multimodal Guidance Network for Missing-Modality Inference in Content Moderation

Zhuokai Zhao, Harish Palani, Tianyi Liu, Lena Evans, Ruth Toner

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICME 2024 Camera Ready. Code is available at https://github.com/zhuokaizhao/multimodal-guidance-network

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22287 2026-03-25 cs.CV cs.AI cs.CL 80%

Founder effects shape the evolutionary dynamics of multimodality in open LLM families

奠基效应塑造了开放大语言模型家族中多模态的进化动态

Manuel Cebrian

机构 * Center for Automation and Robotics, Spanish National Research Council, Madrid, Spain(自动化与机器人中心,西班牙国家研究委员会,西班牙马德里)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究通过分析模型谱系数据,发现多模态能力在开放大语言模型家族中通过罕见的奠基事件进入,并在后代谱系中快速扩展,导致多模态能力的采用动态呈现突变特征。

Comments 7 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11852 2025-10-15 cs.LG 80%

Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection

Saroj Basnet, Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanoji, Marcos Zampieri

机构 * George Mason University(乔治·马歇尔大学) Lancaster University(兰卡斯特大学) University of Surrey(萨里大学)

专题命中 图文多模态 :multimodal(title,abstract)

Comments Accepted to ICDMW 2025 Workshop on Multimodal AI (MMAI). Full workshop info: https://icdmw25mmai.github.io/

Journal ref Proc. IEEE International Conference on Data Mining Workshops (ICDMW 2025), Workshop on Multimodal AI (MMAI 2025), Los Angeles, USA, December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11695 2025-08-25 cs.CV cs.CL cs.MM 80%

Interpreting the linear structure of vision-language model embedding spaces

Isabel Papadimitriou, Huangyuan Su, Thomas Fel, Sham Kakade, Stephanie Gil

机构 * Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University(哈佛大学自然与人工智能研究所) Department of Computer Science, Harvard University(哈佛大学计算机科学系)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.MM

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15392 2025-02-24 cs.AI cs.CL cs.CV 80%

Chitrarth: Bridging Vision and Language for a Billion People

Shaharukh Khan, Ayush Tarun, Abhinav Ravi, Ali Faraz, Akshat Patidar, Praveen Kumar Pokala, Anagha Bhangare, Raja Kolla, Chandra Khatri, Shubham Agarwal

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01278 2023-10-25 cs.CV cs.AI cs.CL 80%

VPGTrans: Transfer Visual Prompt Generator across LLMs

Ao Zhang, Hao Fei, Yuan Yao, Wei Ji, Li Li, Zhiyuan Liu, Tat-Seng Chua

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Website: https://vpgtrans.github.io Code: https://github.com/VPGTrans/VPGTrans NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13437 2023-03-28 cs.CV cs.CL cs.MM 80%

Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning

Yatai Ji, Rongcheng Tu, Jie Jiang, Weijie Kong, Chengfei Cai, Wenzhe Zhao, Hongfa Wang, Yujiu Yang, Wei Liu

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.MM

Comments CVPR 2023 accept

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12737 2022-11-24 cs.CV cs.AI cs.CL cs.LG 80%

RoentGen: Vision-Language Foundation Model for Chest X-ray Generation

Pierre Chambon, Christian Bluethgen, Jean-Benoit Delbrouck, Rogier Van der Sluijs, Małgorzata Połacin, Juan Manuel Zambrano Chaves, Tanishq Mathew Abraham, Shivanshu Purohit, Curtis P. Langlotz, Akshay Chaudhari

专题命中 图文多模态 :multimodal(abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02592 2026-08-25 cs.CV cs.LG 版本更新 79%

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation

H-OPD:置信感知异构多教师多模态在线策略蒸馏

Qixiang Yin, Huanjin Yao, Yuchen Cai, Jianghao Chen, Ziyi Wang, Min Yang, Fei Su, Zhicheng Zhao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ByteDance(字节跳动) USTC(中国科学技术大学) Beijing Key Laboratory of Network System and Network Culture(北京网络系统与网络文化重点实验室) Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism(文化和旅游部互动技术与体验系统重点实验室) Zhongguancun Academy(中关村科学城)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 研究多模态推理的在线策略蒸馏问题,提出H-OPD框架,通过验证异构教师互补性,以令牌级教师仲裁取代任务或样本级路由,结合视觉与文本教师,在多基准测试中性能优越。

Comments EMNLP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15056 2026-08-18 cs.AI 新提交 79%

GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG

GraphLoom:面向多模态KG-RAG的可靠性校准图证据路由

Zafar Ali, Asad Khan, Aalia Malik, Pavlos Kefalas

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Aristotle University(亚里士多德大学) Dashub(达舒布)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 提出GraphLoom框架,通过可靠性校准的图证据路由实现多模态KG-RAG,在ScienceQA等数据集上提升了答案质量与证据忠实度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06223 2026-08-18 cs.CV 79%

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

RedDiffuser:通过强化扩散审计多模态安全故障

Ruofan Wang, Xingjun Ma

机构 * Fudan University(复旦大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 研究多模态系统在有害上下文暴露下的安全审计,提出RedDiffuser框架通过扩散模型生成视觉输入,揭示隐藏的安全漏洞,实验显示VLMs在部分有毒文本与视觉上下文结合时存在广泛安全问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13861 2026-08-17 cs.CV 新提交 79%

XSA-MAD: Cross-modal Semantic Alignment for Morphing Attack Detection

XSA-MAD:用于变形攻击检测的跨模态语义对齐

Jie Jin, Mahiro Tokumasu, Yu Makino, Masakatsu Nishigaki, Tetsushi Ohki

机构 * RIKEN AIP(理化学研究所人工智能项目)

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);分类 cs.CV

AI总结 针对现有图像型变形攻击检测方法泛化性差的问题,提出基于CLIP的多模态框架XSA-MAD,通过对齐图像与语义空间实现对不同变形攻击的检测,在高保真攻击下性能优于现有方法

Comments accepted to ICIP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11013 2026-08-12 cs.CV 新提交 79%

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning

观看合成视频:针对零样本视频字幕生成的视觉合成跨模态表征对齐

Liangyu Fu, Junbo Wang, Yuke Li, Ya Jing, Xuecheng Wu, Zhiyong Wang

机构 * School of Software, Northwestern Polytechnical University(西北工业大学软件学院) School of Information Science and Technology, Beijing University of Technology(北京工业大学信息科学与技术学院) School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) School of Computer Science, The University of Sydney(悉尼大学计算机科学学院)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出WSV零样本视频字幕生成框架,通过文本到视频生成模型、抛光器、提示器弥合跨模态差距,在三个公开数据集上取得了B@4 52、CIDEr 95.7的成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10408 2026-08-12 cs.CL 新提交 79%

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

VisEditBench:视觉语言模型能否基于多模态反馈编辑可视化代码?

Mizanur Rahman, Arshia Azimlu, Shadikur Rahman, Md Tahmid Rahman Laskar, Amran Bhuiyan, Shafiq Joty, Enamul Hoque Prince

机构 * York University(约克大学) Nanyang Technological University(南洋理工大学) Salesforce AI Research(Salesforce人工智能研究)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本研究推出可视化代码编辑基准VisEditBench,评估20种VLMs的编辑能力,提出基于渲染的VisEditAgent框架,显著提升了编辑任务的通过率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09202 2026-08-11 cs.AI 新提交 79%

CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving

CRUISE:面向鲁棒自动驾驶的视觉语言模型引导的不确定性感知跨模态传感器融合

Junyao Wang, Yulin Xu, Yu Li, Pramod Khargonekar, Mohammad Abdullah Al Faruque

机构 * University of California, Irvine(加利福尼亚大学欧文分校)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.AI

AI总结 针对现有不确定性感知融合方法泛化性差的问题,提出CRUISE框架,结合VLM引导的UQ模块与动态自适应机制,实现鲁棒的自动驾驶跨模态传感器融合。

详情

展开后加载摘要…

URL PDF HTML 收藏