arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2508.02762 2025-08-07 cs.LG cs.AI 57%

Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment

Dahun Kim, Anelia Angelova

机构 * Google DeepMind(谷歌DeepMind)

专题命中 图文多模态 :image-text(abstract);分类 cs.AI

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01921 2025-08-05 cs.CV 57%

InspectVLM: Unified in Theory, Unreliable in Practice

Conor Wallace, Isaac Corley, Jonathan Lwowski

机构 * Zeitview

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to 2025 ICCV VISION Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01558 2025-08-05 cs.CV 57%

EvoVLMA: Evolutionary Vision-Language Model Adaptation

Kun Ding, Ying Wang, Shiming Xiang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments This paper has been accepted by ACM Multimedia 2025 (ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00945 2025-08-05 cs.CV 57%

Optimizing Vision-Language Consistency via Cross-Layer Regional Attention Alignment

Yifan Wang, Hongfeng Ai, Quangao Liu, Maowei Jiang, Ruiyuan Kang, Ruiqi Li, Jiahua Dong, Mengting Xiao, Cheng Jiang, Chenzhong Li

机构 * School of Medicine, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)医学院) Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Wave and Machine Intelligence Department, Technology Innovation Institute(技术创新研究院波浪与机器智能部门) University of the Chinese Academy of Sciences(中国科学院大学) McGill University(麦吉尔大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12441 2025-08-05 cs.CV cs.LG 57%

Describe Anything Model for Visual Question Answering on Text-rich Images

Yen-Linh Vu, Dinh-Thang Duong, Truong-Binh Duong, Anh-Khoi Nguyen, Thanh-Huy Nguyen, Le Thien Phuc Nguyen, Jianhua Xing, Xingjian Li, Tianyang Wang, Ulas Bagci, Min Xu

机构 * AI VIETNAM Lab(AI越南实验室) Carnegie Mellon University(卡内基梅隆大学) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) University of Pittsburgh(匹兹堡大学) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) Northwestern University(西北大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 11 pages, 5 figures. Accepted to VisionDocs @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18645 2025-08-05 cs.CV 57%

Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings

Azim Ospanov, Mohammad Jalali, Farzan Farnia

机构 * The Chinese University of Hong Kong, Department of Computer Science & Engineering(香港中文大学计算机科学与工程系)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00356 2025-08-04 cs.CV cs.MA 57%

Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning

Angelos Vlachos, Giorgos Filandrianos, Maria Lymperaiou, Nikolaos Spanos, Ilias Mitsouras, Vasileios Karampinis, Athanasios Voulodimos

机构 * Artificial Intelligence and Learning Systems Laboratory, National Technical University of Athens(人工智能与学习系统实验室,国家技术大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19480 2025-08-01 cs.CV 57%

GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers

Shijie Ma, Yuying Ge, Teng Wang, Yuxin Guo, Yixiao Ge, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Institute of Automation, CAS(中国科学院自动化研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025. Project released at: https://mashijie1028.github.io/GenHancer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23362 2025-08-01 cs.CV 57%

Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers

Ji Ma, Wei Suo, Peng Wang, Yanning Zhang

机构 * Northwestern Polytechnical University(西北工业大学) National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted By ACM MM 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23226 2025-08-01 cs.CV 57%

Toward Safe, Trustworthy and Realistic Augmented Reality User Experience

Yanming Xiu

机构 * Department of Electrical and Computer Engineering, Duke University(电子工程系,杜克大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 2 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22346 2025-07-31 cs.CV 57%

DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception

Pei Deng, Wenqian Zhou, Hanlin Wu

机构 * School of Information Science and Technology, Beijing Foreign Studies University(信息科学与技术学院,北京外国语大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 5 figures. Submitted to IEEE Transactions on Geoscience and Remote Sensing (TGRS). Code and dataset are available at https://github.com/hanlinwu/DeltaVLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22003 2025-07-31 cs.CV 57%

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs

Ziyun Dai, Xiaoqiang Li, Shaohua Zhang, Yuanchen Wu, Jide Li

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM25

Journal ref 33rd ACM International Conference on Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12447 2025-07-31 cs.CV 57%

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification

Nathanael L. Baisa, Babu Pallam, Amudhavel Jayavel

机构 * School of Computer Science(计算机科学学院) Informatics De Montfort University(信息学德蒙特福特大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22101 2025-07-31 cs.CV 57%

AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock

Umair Nawaz, Muhammad Zaigham Zaheer, Fahad Shahbaz Khan, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer

机构 * MBZ University of AI(人工智能大学) CECS, Australian National University(计算机科学与工程系,澳大利亚国立大学) Computer Vision Laboratory, Linköping University(链接öping大学计算机视觉实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21794 2025-07-30 cs.CV 57%

Distribution-Based Masked Medical Vision-Language Model Using Structured Reports

Shreyank N Gowda, Ruichi Zhang, Xiao Gu, Ying Weng, Lu Yang

机构 * School of Computer Science, University of Nottingham(计算机科学学院,诺丁汉大学) Department of Computer Science and Technology, School of Informatics, Xiamen University(计算机科学与技术系,信息学院,厦门大学) CHI Lab, University of Oxford(CHI实验室,牛津大学) School of Computer Science, University of Nottingham Ningbo China(计算机科学学院,宁波大学中国)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted in MICCAI-W 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21291 2025-07-30 cs.CV 57%

Fairness and Robustness of CLIP-Based Models for Chest X-rays

Théo Sourget, David Restrepo, Céline Hudelot, Enzo Ferrante, Stergios Christodoulidis, Maria Vakalopoulou

机构 * MICS, CentraleSupélec - Université Paris-Saclay(MICS,中央超导学院——巴黎萨克雷大学) CONICET, Universidad de Buenos Aires(CONICET,布宜诺斯艾利斯大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted for publication at the FAIMI MICCAI workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21165 2025-07-30 eess.IV cs.CV 57%

Querying GI Endoscopy Images: A VQA Approach

Gaurav Parajuli

机构 * Johannes Kepler University Linz(约翰·凯撒大学林茨)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19651 2025-07-30 cs.CV cs.LG cs.PF 57%

PEVLM: Parallel Encoding for Vision-Language Models

Letian Kang, Shixian Luo, Yiqiang Li, Yuxin Yin, Shenxuan Zhou, Xiaoyang Yu, Jin Yang, Yong Wu

机构 * Li Auto Inc.(利自动公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20397 2025-07-29 cs.CV 57%

VESPA: Towards un(Human)supervised Open-World Pointcloud Labeling for Autonomous Driving

Levente Tempfli, Esteban Rivera, Markus Lienkamp

机构 * Technical University of Munich(慕尼黑技术大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20188 2025-07-29 cs.CV 57%

SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection

Mohammed-En-Nadhir Zighem, Abdenour Hadid

机构 * Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi, UAE(索邦人工智能中心,阿布扎比分校,阿联酋)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19875 2025-07-29 cs.CV 57%

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking

X. Feng, S. Hu, X. Li, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang

机构 * School of Artificial Intelligence, UCAS(人工智能学院,中国科学院大学) The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, CASIA(复杂系统认知与决策智能重点实验室,中国科学院自动化所) School of Physical and Mathematical Sciences, NTU(物理与数学科学学院,国立新加坡大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV2025 Highlight ~

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15480 2025-07-29 cs.CV 57%

One Last Attention for Your Vision-Language Model

Liang Chen, Ghazi Shazan Ahmad, Tianjun Yao, Lingqiao Liu, Zhiqiang Shen

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17762 2025-07-29 cs.CV 57%

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Rongchang Xie, Chen Du, Ping Song, Chang Liu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10341 2025-07-29 cs.RO cs.AI cs.LG 57%

Affordance-Guided Reinforcement Learning via Visual Prompting

Olivia Y. Lee, Annie Xie, Kuan Fang, Karl Pertsch, Chelsea Finn

机构 * Stanford University(斯坦福大学) Cornell University(康奈尔大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

Comments 8 pages, 6 figures. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18433 2025-07-25 eess.IV cs.CV 57%

DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis

Minxi Ouyang, Lianghui Zhu, Yaqing Bao, Qiang Huang, Jingli Ouyang, Tian Guan, Xitong Ling, Jiawen Li, Song Duan, Wenbin Dai, Li Zheng, Xuemei Zhang, Yonghong He

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Department of Pathology, Liuzhou People’s Hospital Affiliated to Guangxi Medical University(广西医科大学柳州市人民医院病理科) Department of Immunology, College of Basic Medical Sciences, China Medical University(中国医科大学基础医学学院免疫科) Greater Bay Area Center for Medical Device Evaluation and Inspection.NMPA(粤港澳大湾区医疗器械评价和检验中心.NMPA) Shenzhen Shengqiang Technology Co., Ltd.(深圳盛强科技有限公司) Department of Pathology, Chongqing University Affiliated Three Gorges Hospital(重庆大学附属第三人民医院病理科)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17722 2025-07-24 cs.CV 57%

BetterCheck: Towards Safeguarding VLMs for Automotive Perception Systems

Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu, Christian Berger

机构 * University of Gothenburg(哥德堡大学) Chalmers University of Technology(查尔姆斯理工大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in The IEEE International Conference on Intelligent Transportation Systems (ITSC)2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17239 2025-07-24 cs.CV 57%

MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training

Lei Zhu, Jun Zhou, Rick Siow Mong Goh, Yong Liu

机构 * Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(高性能计算研究所(IHPC)、科技研究局(A*STAR))

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to MedAGI 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09184 2025-07-24 cs.CV 57%

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models

Qiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing, Xiaosong Yuan, Feilong Tang, Sinan Fan, Xuhang Chen, Xuyao Zhang, Dahan Wang

机构 * FKLPRIU, Xiamen University of Technology, China(福克斯理工国际大学,厦门理工学院,中国) Shanghai Jiao Tong University, China(上海交通大学,中国) Nanyang Technological University, Singapore(南洋理工大学,新加坡) Jilin University, China(吉林大学,中国) Monash University, Australia(墨尔本大学,澳大利亚) Zhejiang University, China(浙江大学,中国) Huizhou University, China(惠州大学,中国) Chinese Academy of Sciences, China(中国科学院,中国)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09815 2025-07-23 cs.CV 57%

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding

Younggun Kim, Ahmed S. Abdelrahman, Mohamed Abdel-Aty

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 22 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10440 2025-07-22 cs.CV 57%

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Guowei Xu, Peng Jin, Ziang Wu, Hao Li, Yibing Song, Lichao Sun, Li Yuan

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 17 pages, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏