arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46237 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4672 篇

2508.10339 2025-08-15 cs.CV cs.LG 74%

Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models

Andrew Bai, Justin Cui, Ruochen Wang, Cho-Jui Hsieh

机构 * Department of Computer Science University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15621 2025-08-01 cs.CV cs.AI cs.CL cs.MM 74%

LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Federico Cocchi, Nicholas Moratelli, Davide Caffagni, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学) IIT-CNR(意大利国家研究 council(IIT))

专题命中 图文多模态 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)

Comments ICCV 2025 Workshop on What is Next in Multimodal Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19370 2025-07-28 cs.CV 74%

BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving

Felix Brandstaetter, Erik Schuetz, Katharina Winter, Fabian Flohr

机构 * Intelligent Vehicles Lab (IVL) Munich University of Applied Sciences(智能车辆实验室(IVL)慕尼黑应用科学大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14953 2025-07-18 cs.CV 74%

Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching

Yang Liu, Wentao Feng, Zhuoyao Liu, Shudong Huang, Jiancheng Lv

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Engineering Research Center of Machine Learning and Industry Intelligence(机器学习与工业智能工程研究中心)

专题命中 图文多模态 :image-text(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12236 2025-07-17 cs.CV 74%

Generate to Ground: Multimodal Text Conditioning Boosts Phrase Grounding in Medical Vision-Language Models

Felix Nützel, Mischa Dombrowski, Bernhard Kainz

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(弗赖堡-亚历山大大学埃尔兰根-纽伦堡) Imperial College London(伦敦帝国理工学院)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments 20 pages, 6 figures. To appear in Proc. MIDL 2025 (PMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18491 2025-06-12 cs.CL 74%

MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering

Shuo Yang, Siwen Luo, Soyeon Caren Han, Eduard Hovy

机构 * The University of Melbourne(墨尔本大学) The University of Western Australia(西澳大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CL

Comments Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05441 2025-06-09 eess.IV cs.CV cs.LG 74%

Deep histological synthesis from mass spectrometry imaging for multimodal registration

Kimberley M. Bird, Xujiong Ye, Alan M. Race, James M. Brown

机构 * University of Lincoln(林肯大学) University of Exeter(埃克塞特大学) AstraZeneca Computational Pathology GmbH(阿斯利康计算病理学 GmbH)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments Medical Image Understanding and Analysis (MIUA) 2025 Extended Abstract Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20937 2025-05-28 cs.CL 74%

On VLMs for Diverse Tasks in Multimodal Meme Classification

Deepesh Gavit, Debajyoti Mazumder, Samiran Das, Jasabanta Patro

机构 * Peng Wang and Shuai Bai and Sinan Tan and Shijie Wang and Zhihao Fan and Jinze Bai and Keqin Chen and Xuejing Liu and Jialin Wang and Wenbin Ge and Yang Fan and Kai Dang and Mengfei Du and Xuancheng Ren and Rui Men and Dayiheng Liu and Chang Zhou and Jingren Zhou and Junyang Lin(研究人员)

专题命中 图文多模态 :multimodal(title);分类 cs.CL

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11576 2025-03-17 cs.CV 74%

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Ahmed Nassar, Andres Marafioti, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, Lucas Morin, Rafael Teixeira de Lima, Yusik Kim, A. Said Gurbuz, Michele Dolfi, Miquel Farré, Peter W. J. Staar

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments 24 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17251 2024-12-24 cs.CV cs.LG eess.IV 74%

GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning

Teja Krishna Cherukuri, Nagur Shareef Shaik, Jyostna Devi Bodapati, Dong Hye Ye

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments This paper has been accepted for presentation at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01725 2024-12-03 cs.CV 74%

Attacks on multimodal models

Viacheslav Iablochnikov, Alexander Rogachev

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments 19 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19296 2024-11-19 cs.AI 74%

Multi-Modal CLIP-Informed Protein Editing

Mingze Yin, Hanjing Zhou, Yiheng Zhu, Miao Lin, Yixuan Wu, Jialu Wu, Hongxia Xu, Chang-Yu Hsieh, Tingjun Hou, Jintai Chen, Jian Wu

专题命中 图文多模态 :multi-modal(title);分类 cs.AI

Comments 13 pages, 7 figures, 5 tables

Journal ref Health Data Science, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07241 2024-11-05 cs.CV cs.LG 74%

Calibrating Multi-modal Representations: A Pursuit of Group Robustness without Annotations

Chenyu You, Yifei Min, Weicheng Dai, Jasjeet S. Sekhon, Lawrence Staib, James S. Duncan

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12736 2024-07-18 cs.CV 74%

Towards Multimodal In-Context Learning for Vision & Language Models

Sivan Doveh, Shaked Perek, M. Jehanzeb Mirza, Wei Lin, Amit Alfassy, Assaf Arbelle, Shimon Ullman, Leonid Karlinsky

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06586 2024-05-13 cs.CV 74%

Enhancing Weakly Supervised Semantic Segmentation with Multi-modal Foundation Models: An End-to-End Approach

Elham Ravanbakhsh, Cheng Niu, Yongqing Liang, J. Ramanujam, Xin Li

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00119 2023-12-05 cs.CV 74%

Fewshot learning on global multimodal embeddings for earth observation tasks

Matt Allen, Francisco Dorr, Joseph A. Gallego-Mejia, Laura Martínez-Ferrer, Anna Jungbluth, Freddie Kalaitzis, Raúl Ramos-Pollán

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments 9 pages, 6 figures, presented on NeurIPS workshop on Robustness of Few-shot and Zero-shot Learning in Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16529 2023-06-30 cs.IR cs.AI cs.DL 74%

Multimodal Search on Iconclass using Vision-Language Pre-Trained Models

Cristian Santini, Etienne Posthumus, Mary Ann Tan, Oleksandra Bruns, Tabea Tietz, Harald Sack

专题命中 图文多模态 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02828 2023-04-07 cs.CV cs.CY 74%

Uncurated Image-Text Datasets: Shedding Light on Demographic Bias

Noa Garcia, Yusuke Hirota, Yankun Wu, Yuta Nakashima

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17169 2023-03-31 cs.CV 74%

Task-Oriented Multi-Modal Mutual Leaning for Vision-Language Models

Sifan Long, Zhen Zhao, Junkun Yuan, Zichang Tan, Jiangjiang Liu, Luping Zhou, Shengsheng Wang, Jingdong Wang

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13333 2022-09-07 cs.CV cs.GR cs.LG 74%

CLIP-Mesh: Generating textured meshes from text using pretrained image-text models

Nasir Mohammad Khalid, Tianhao Xie, Eugene Belilovsky, Tiberiu Popa

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments 8 pages, 8 figures, Accepted at SIGGRAPH ASIA 2022, Project Page at https://www.nasir.lol/clipmesh

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.11848 2021-09-27 cs.CV 74%

How to find a good image-text embedding for remote sensing visual question answering?

Christel Chappuis, Sylvain Lobry, Benjamin Kellenberger, Bertrand Le Saux, Devis Tuia

专题命中 图文多模态 :image-text(title);分类 cs.CV

Comments 10 pages, 4 figures, presented in the MACLEAN workshop during ECML PKDD 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08106 2021-05-19 cs.CL 74%

Multi-Modal Image Captioning for the Visually Impaired

Hiba Ahsan, Nikita Bhalla, Daivat Bhatt, Kaivankumar Shah

专题命中 图文多模态 :multi-modal(title);分类 cs.CL

Comments 8 pages, 2 figures, 2 tables, accepted to NAACL-HLT SRW 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.03315 2020-06-08 cs.CV cs.LG eess.IV 74%

Multi-modal Feature Fusion with Feature Attention for VATEX Captioning Challenge 2020

Ke Lin, Zhuoxin Gan, Liwei Wang

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11416 2019-09-26 cs.MM 74%

Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching

Chunxiao Liu, Zhendong Mao, An-An Liu, Tianzhu Zhang, Bin Wang, Yongdong Zhang

专题命中 图文多模态 :image-text(title);分类 cs.MM

Comments Accepted by ACMMM2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.09317 2019-08-27 cs.CV 74%

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings

Iro Laina, Christian Rupprecht, Nassir Navab

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments ICCV 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.01919 2019-08-07 cs.CV 74%

Image Captioning with Clause-Focused Metrics in a Multi-Modal Setting for Marketing

Philipp Harzig, Dan Zecha, Rainer Lienhart, Carolin Kaiser, René Schallner

专题命中 图文多模态 :multi-modal(title);分类 cs.CV

Comments 6 pages, accepted at MIPR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.01958 2019-08-07 cs.CV 74%

Multimodal Image Captioning for Marketing Analysis

Philipp Harzig, Stephan Brehm, Rainer Lienhart, Carolin Kaiser, René Schallner

专题命中 图文多模态 :multimodal(title);分类 cs.CV

Comments 4 pages, 1 figure, accepted at MIPR2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1512.04701 2015-12-16 cs.IR cs.CL cs.SI 74%

Joint Image-Text News Topic Detection and Tracking with And-Or Graph Representation

Weixin Li, Jungseock Joo, Hang Qi, Song-Chun Zhu

专题命中 图文多模态 :image-text(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20414 2026-08-24 cs.AI cs.CV 新提交 73%

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

StateSight:评测视觉语言模型中的潜在空间状态重建能力

Michelle Lin

机构 * Thomas Jefferson High School for Science and Technology(托马斯·杰斐逊科学技术高中)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本研究推出StateSight基准及配套数据集StateSight-Steps,评估视觉语言模型的潜在空间状态重建能力,发现GPT-5.5、Claude Sonnet 5的表现均逊于人类基线,格式正确的响应可能掩盖空间结构恢复失败的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16841 2026-08-20 cs.CV cs.MM 版本更新 73%

Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment

回答前看清楚:通过显著性驱动的感知重新对齐减轻LVLMs中的幻觉

Pengxu Chen, Yao Zhu, Guangming Zhu, Jun Sheng, Jincai Huang, Xiangyang Ji, Liang Zhang

机构 * Xidian University(西安电子科技大学) Tsinghua University(清华大学) Shanghai Road Transport Development Center(上海市道路运输发展中心) Hunan Institute of Advanced Technology(湖南先进技术研究院)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

AI总结 研究针对LVLMs易产生幻觉问题,提出无需训练的SDPR框架,通过显著性驱动注意力重新分配、缓存对齐及先验约束对比解码,整体对齐视觉意识,在多基准测试中优于现有方法,无需额外训练且开销小。

Comments Accepted by ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏