arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4682 篇

2506.07399 2025-06-10 cs.CV cs.AI 84%

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems

Peiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai, Huili Wang, Yufei Sun, Xintian Li, Shangguang Wang, Yongfeng Huang, Tao Qi

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07227 2025-06-10 cs.CV cs.CL 84%

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

Tianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun, Jiayi Song, Junlin Han, Zichen Liu, Conghui He, Wentao Zhang, Binhang Yuan

机构 * HKUST(香港科技大学) Shanghai AI Lab(上海人工智能实验室) Peking University(北京大学) HKUST(GZ)(香港科技大学(广州)) Oxford University(牛津大学) Imperial College London(伦敦帝国理工学院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15239 2025-06-10 cs.CV cs.AI cs.IR 84%

Benchmark Granularity and Model Robustness for Image-Text Retrieval

Mariya Hendriksen, Shuo Zhang, Ridho Reinanda, Mohamed Yahya, Edgar Meij, Maarten de Rijke

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments accepted at SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04277 2025-06-06 cs.CV cs.AI 84%

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Yi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li, Licheng Tang, Yangguang Ji, Chong Wu, Jay Wu, Wenbo Zhu

专题命中 图文多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted as ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11784 2025-06-05 cs.AI cs.CV cs.LG 84%

Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development

Daoyuan Chen, Haibin Wang, Yilun Huang, Ce Ge, Yaliang Li, Bolin Ding, Jingren Zhou

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICML 2025 (Spotlight). 33 pages, 16 tables, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14035 2025-05-21 cs.MM cs.CL 84%

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs

Shiyao Cui, Qinglin Zhang, Xuan Ouyang, Renmiao Chen, Zhexin Zhang, Yida Lu, Hongning Wang, Han Qiu, Minlie Huang

机构 * The Conversational AI (CoAI) group, DCST, Tsinghua University China(清华大学人工智能对话组,国防科技大学,清华大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00958 2025-05-14 cs.CV cs.CL cs.LG 84%

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

Wenqi Zhang, Hang Zhang, Xin Li, Jiashuo Sun, Yongliang Shen, Weiming Lu, Deli Zhao, Yueting Zhuang, Lidong Bing

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) DAMO Academy, Alibaba Group(阿里巴巴集团达摩院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06152 2025-05-12 cs.CV cs.AI 84%

MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks

Wenqi Zeng, Yuqi Sun, Chenxi Ma, Weimin Tan, Bo Yan

机构 * Shanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University(上海智能信息处理关键实验室,计算机科学学院,复旦大学)

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08802 2025-04-25 cs.CL cs.CV cs.IR 84%

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Andreas Koukounas, Georgios Mastrapas, Sedigheh Eslami, Bo Wang, Mohammad Kalim Akram, Michael Günther, Isabelle Mohr, Saba Sturua, Nan Wang, Han Xiao

机构 * Jina AI GmbH(Jina AI公司)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 30 pages, 1-10 main paper, 10-12 refs, 12-30 benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12902 2025-04-16 cs.AI cs.CL cs.LG 84%

IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities

Bin Wang, Chunyu Xie, Dawei Leng, Yuhui Yin

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10049 2025-04-15 cs.CV cs.CL 84%

Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure

Théo Gigant, Camille Guinaudeau, Frédéric Dufaux

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08269 2025-04-14 cs.CV cs.CL 84%

VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering

Qi Zhi Lim, Chin Poo Lee, Kian Ming Lim, Kalaiarasi Sonai Muthu Anbananthen

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07336 2025-04-11 cs.CV cs.AI 84%

Zeus: Zero-shot LLM Instruction for Union Segmentation in Multimodal Medical Imaging

Siyuan Dai, Kai Ye, Guodong Liu, Haoteng Tang, Liang Zhan

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 4 figures, In Press by a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21839 2025-03-31 cs.CV cs.AI cs.LG 84%

M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?

Haolong Yan, Kaijun Tan, Yeqing Shen, Xin Huang, Zheng Ge, Xiangyu Zhang, Si Li, Daxin Jiang

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09220 2025-03-26 cs.CL cs.AI 84%

LingYi: Medical Conversational Question Answering System based on Multi-modal Knowledge Graphs

Fei Xia, Bin Li, Yixuan Weng, Shizhu He, Kang Liu, Bin Sun, Shutao Li, Jun Zhao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09091 2025-03-21 cs.CV cs.AI 84%

Multi-Modal Foundation Models for Computational Pathology: A Survey

Dong Li, Guihong Wan, Xintao Wu, Xinyu Wu, Xiaohui Chen, Yi He, Christine G. Lian, Peter K. Sorger, Yevgeniy R. Semenov, Chen Zhao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04543 2025-03-07 cs.CL cs.AI 84%

Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model

Wenke Huang, Jian Liang, Xianda Guo, Yiyang Fang, Guancheng Wan, Xuankun Rong, Chi Wen, Zekun Shi, Qingyun Li, Didi Zhu, Yanbiao Ma, Ke Liang, Bin Yang, He Li, Jiawei Shao, Mang Ye, Bo Du

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20277 2025-02-28 cs.CV cs.AI 84%

Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions

Palawat Busaranuvong, Emmanuel Agu, Reza Saadati Fard, Deepak Kumar, Shefalika Gautam, Bengisu Tulu, Diane Strong

专题命中 图文多模态 :multi-modal(title);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01703 2025-02-03 cs.CL cs.AI cs.LG 84%

UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models

Sejoon Oh, Yiqiao Jin, Megha Sharma, Donghyun Kim, Eric Ma, Gaurav Verma, Srijan Kumar

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10489 2024-12-25 cs.CV cs.AI eess.SP 84%

CognitionCapturer: Decoding Visual Stimuli From Human EEG Signal With Multimodal Information

Kaifan Zhang, Lihuo He, Xin Jiang, Wen Lu, Di Wang, Xinbo Gao

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07292 2024-12-11 cs.MM cs.CL 84%

Multimodal Sentiment Analysis Based on Causal Reasoning

Fuhai Chen, Pengpeng Huang, Xuri Ge, Jie Huang, Zishuo Bao

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07112 2024-12-11 cs.CV cs.CL 84%

Maya: An Instruction Finetuned Multilingual Multimodal Model

Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, S M Iftekhar Uddin, Shayekh Bin Islam, Roshan Santhosh, Snegha A, Drishti Sharma, Chen Liu, Isha Chaturvedi, Genta Indra Winata, Ashvanth. S, Snehanshu Mukherjee, Alham Fikri Aji

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05939 2024-12-10 cs.CV cs.CL cs.LG 84%

Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models

Xiao Xu, Tianhao Niu, Yuxi Xie, Libo Qin, Wanxiang Che, Min-Yen Kan

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments A manuscript that should have been Arxived in May :)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21414 2024-10-30 cs.CL cs.AI 84%

CT2C-QA: Multimodal Question Answering over Chinese Text, Table and Chart

Bowen Zhao, Tianhao Cheng, Yuejie Zhang, Ying Cheng, Rui Feng, Xiaobo Zhang

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02026 2024-10-01 cs.CV cs.AI 84%

Multimodal-Enhanced Objectness Learner for Corner Case Detection in Autonomous Driving

Lixing Xiao, Ruixiao Shi, Xiaoyang Tang, Yi Zhou

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to 2024 IEEE International Conference on Image Processing (ICIP) as oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19404 2024-09-23 cs.CV cs.CL 84%

EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning

Junzhe Zhang, Huixuan Zhang, Xunjian Yin, Xiaojun Wan

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06731 2024-08-29 cs.CV cs.AI 84%

Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator

Henry Hengyuan Zhao, Pan Zhou, Mike Zheng Shou

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13461 2024-08-27 cs.CV cs.AI 84%

Probing the Robustness of Vision-Language Pretrained Models: A Multimodal Adversarial Attack Approach

Jiwei Guan, Tianyu Ding, Longbing Cao, Lei Pan, Chen Wang, Xi Zheng

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03149 2024-08-07 cs.CV cs.CL 84%

Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization

Yanghai Zhang, Ye Liu, Shiwei Wu, Kai Zhang, Xukai Liu, Qi Liu, Enhong Chen

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments In ACL-Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13488 2024-07-19 cs.CV cs.MM 84%

Similarity over Factuality: Are we making progress on multimodal out-of-context misinformation detection?

Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis C. Petrantonakis

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏