arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4657 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4657 篇

2507.18433 2025-07-25 eess.IV cs.CV 57%

DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis

Minxi Ouyang, Lianghui Zhu, Yaqing Bao, Qiang Huang, Jingli Ouyang, Tian Guan, Xitong Ling, Jiawen Li, Song Duan, Wenbin Dai, Li Zheng, Xuemei Zhang, Yonghong He

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Department of Pathology, Liuzhou People’s Hospital Affiliated to Guangxi Medical University(广西医科大学柳州市人民医院病理科) Department of Immunology, College of Basic Medical Sciences, China Medical University(中国医科大学基础医学学院免疫科) Greater Bay Area Center for Medical Device Evaluation and Inspection.NMPA(粤港澳大湾区医疗器械评价和检验中心.NMPA) Shenzhen Shengqiang Technology Co., Ltd.(深圳盛强科技有限公司) Department of Pathology, Chongqing University Affiliated Three Gorges Hospital(重庆大学附属第三人民医院病理科)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17722 2025-07-24 cs.CV 57%

BetterCheck: Towards Safeguarding VLMs for Automotive Perception Systems

Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu, Christian Berger

机构 * University of Gothenburg(哥德堡大学) Chalmers University of Technology(查尔姆斯理工大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in The IEEE International Conference on Intelligent Transportation Systems (ITSC)2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17239 2025-07-24 cs.CV 57%

MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training

Lei Zhu, Jun Zhou, Rick Siow Mong Goh, Yong Liu

机构 * Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(高性能计算研究所(IHPC)、科技研究局(A*STAR))

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to MedAGI 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09184 2025-07-24 cs.CV 57%

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models

Qiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing, Xiaosong Yuan, Feilong Tang, Sinan Fan, Xuhang Chen, Xuyao Zhang, Dahan Wang

机构 * FKLPRIU, Xiamen University of Technology, China(福克斯理工国际大学,厦门理工学院,中国) Shanghai Jiao Tong University, China(上海交通大学,中国) Nanyang Technological University, Singapore(南洋理工大学,新加坡) Jilin University, China(吉林大学,中国) Monash University, Australia(墨尔本大学,澳大利亚) Zhejiang University, China(浙江大学,中国) Huizhou University, China(惠州大学,中国) Chinese Academy of Sciences, China(中国科学院,中国)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09815 2025-07-23 cs.CV 57%

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding

Younggun Kim, Ahmed S. Abdelrahman, Mohamed Abdel-Aty

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 22 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10440 2025-07-22 cs.CV 57%

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Guowei Xu, Peng Jin, Ziang Wu, Hao Li, Yibing Song, Lichao Sun, Li Yuan

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 17 pages, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11055 2025-07-22 cs.CV 57%

Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic Approximation

Shuchang Ye, Usman Naseem, Mingyuan Meng, Jinman Kim

机构 * The University of Sydney(悉尼大学) Macquarie University(麦觉瑞大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09222 2025-07-22 cs.CV cs.LG 57%

Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift

Behraj Khan, Tahir Qasim Syed, Nouman M. Durrani, Bilal Naseem, Shabir Ahmad, Rizwan Qureshi

机构 * Institute of Business Administration Karachi(Karachi商业管理学院) National University of Computer and Emerging Sciences(国家计算机与新兴科学大学) CAIMI Pvt Ltd(CAIMI私营有限公司) Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,佛罗里达大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10283 2025-07-22 cs.RO cs.AI cs.SY eess.IV eess.SY 57%

ASMA: An Adaptive Safety Margin Algorithm for Vision-Language Drone Navigation via Scene-Aware Control Barrier Functions

Sourav Sanyal, Kaushik Roy

机构 * School of Electrical and Computer Engineering, Purdue University(电气与计算机工程学院,普渡大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11200 2025-07-21 cs.CV 57%

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

Che Liu, Jiazhen Pan, Weixiang Shen, Wenjia Bai, Daniel Rueckert, Rossella Arcucci

机构 * Imperial College London, UK(伦敦帝国学院) Technical University of Munich, Germany(慕尼黑技术大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13359 2025-07-21 cs.CV 57%

Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives

Yang Zhou, Junjie Li, CongYang Ou, Dawei Yan, Haokui Zhang, Xizhe Xue

机构 * Northwestern Polytechnical University(北华理工大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments 27 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13113 2025-07-18 cs.CV 57%

Leveraging Language Prior for Infrared Small Target Detection

Pranav Singh, Pravendra Singh

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Roorkee(计算机科学与工程系,印度理工学院Roorkee)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12755 2025-07-18 cs.CV cs.LG 57%

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation

Yanchen Guan, Haicheng Liao, Chengyue Wang, Bonan Wang, Jiaxun Zhang, Jia Hu, Zhenning Li

机构 * State Key Laboratory of Internet of Things for Smart City(物联网智能城市国家重点实验室) University of Macau(澳门大学) Department of Civil Engineering(土木工程系) Department of Computer and Information Science(计算机与信息科学系) College of Transportation Engineering(交通工程学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03770 2025-07-18 cs.CR cs.AI 57%

JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model

Yi Nian, Shenzhe Zhu, Yuehan Qin, Li Li, Ziyi Wang, Chaowei Xiao, Yue Zhao

机构 * University of Southern California(南加州大学) University of Toronto(多伦多大学) University of Maryland(马里兰大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19697 2025-07-18 cs.CV 57%

Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion

Yuan Bian, Min Liu, Yunqi Yi, Xueping Wang, Yaonan Wang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) College of Information Science and Engineering, Hunan Normal University(湖南师范大学信息科学与工程学院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11102 2025-07-16 cs.CV 57%

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

Jie Yang, Wang Zeng, Sheng Jin, Lumin Xu, Wentao Liu, Chen Qian, Zhen Li, Ruimao Zhang

机构 * Sun Yat-sen University, Shenzhen(中山大学深圳校区) Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区) SenseTime Research(商汤科技研究院) Chinese University of Hong Kong(香港中文大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Extended Version of KptLLM. arXiv admin note: text overlap with arXiv:2411.01846

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05464 2025-07-16 cs.CL 57%

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Shiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu, Siyang Gao, Miao Xiong, Manling Li, Junxian He

机构 * City University of Hong Kong(香港城市大学) Hong Kong University of Science(香港科学大学) National University of Singapore(新加坡国立大学) Northwestern University(西北大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments ICML 2025. Camera-ready version updated. Our code is publicly available at https://github.com/shiqichen17/VLM_Merging

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01069 2025-07-15 cs.CV 57%

UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment

Hantao Zhou, Longxiang Tang, Rui Yang, Guanyi Qin, Yan Zhang, Yutao Li, Xiu Li, Runze Hu, Guangtao Zhai

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Media Analytics and Computing Lab, Department of Artificial Intelligence, School of Informatics, Xiamen University(媒体分析与计算实验室,人工智能系,厦门大学) School of Computer Science and Technology, Ocean University of China(计算机科学与技术学院,中国海洋大学) Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(图像通信与信息处理研究所,上海交通大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09615 2025-07-15 cs.CV 57%

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score

Eman Ali, Sathira Silva, Chetan Arora, Muhammad Haris Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫莫德·本·扎耶德人工智能大学) IIT Delhi(德里印度理工学院) Alexandria University(亚历山大大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09008 2025-07-15 cs.CV 57%

VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels

Xiwei Xuan, Xiaoqi Wang, Wenbin He, Jorge Piazentin Ono, Liang Gou, Kwan-Liu Ma, Liu Ren

机构 * Department of Computer Science, University of California, Davis, CA, USA(加州大学戴维斯分校计算机科学系) Bosch Center for Artificial Intelligence (BCAI), Bosch Research North America(博世人工智能中心(BCAI)、博世北美研究部) Splunk Technology, San Jose, CA, USA(Splunk技术公司)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments IEEE Transactions on Visualization and Computer Graphics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08982 2025-07-15 eess.IV cs.CV cs.LG 57%

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models

Hanene F. Z. Brachemi Meftah, Wassim Hamidouche, Sid Ahmed Fezza, Olivier Déforges

机构 * Univ. Rennes, INSA Rennes, CNRS, IETR - UMR 6164(里昂大学、里昂国家理工学院、国家科学研究中心、IETR - UMR 6164)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07104 2025-07-14 cs.CV 57%

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models

Tiezheng Zhang, Yitong Li, Yu-cheng Chou, Jieneng Chen, Alan Yuille, Chen Wei, Junfei Xiao

机构 * Johns Hopkins University(约翰霍普金斯大学) Tsinghua University(清华大学) Rice University(Rice 大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Project Page: https://lambert-x.github.io/Vision-Language-Vision/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07562 2025-07-11 cs.CL 57%

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs

Jierun Chen, Tiezheng Yu, Haoli Bai, Lewei Yao, Jiannan Wu, Kaican Li, Fei Mi, Chaofan Tao, Lei Zhu, Manyi Zhang, Xiaohui Li, Lu Hou, Lifeng Shang, Qun Liu

机构 * Huawei Technologies(华为技术有限公司) HKUST(香港科技大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07148 2025-07-11 cs.CV cs.LG 57%

Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey

Getamesay Haile Dagnaw, Yanming Zhu, Muhammad Hassan Maqsood, Wencheng Yang, Xingshuai Dong, Xuefei Yin, Alan Wee-Chung Liew

机构 * Griffith University(格里菲斯大学) University of Southern Queensland(南部昆士兰大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17670 2025-07-09 cs.LG cs.AI 57%

Towards General Continuous Memory for Vision-Language Models

Wenyi Wu, Zixuan Song, Kun Zhou, Yifei Shao, Zhiting Hu, Biwei Huang

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04699 2025-07-08 cs.CV 57%

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

Zexi Jia, Chuanwei Huang, Hongyan Fei, Yeshuang Zhu, Zhiqiang Yuan, Ying Deng, Jiapei Zhang, Jinchao Zhang, Jie Zhou

机构 * WeChat AI, Tencent Inc, China(腾讯公司)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05163 2025-07-08 cs.CV cs.LG 57%

Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models

Aishwarya Venkataramanan, Paul Bodesheim, Joachim Denzler

机构 * Computer Vision Group, Friedrich Schiller University Jena(计算机视觉组,费迪里奇·施勒尔大学耶纳)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments UAI 2025, 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20612 2025-07-08 cs.CV 57%

IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware Prompting

Hao Fu, Hanbin Zhao, Jiahua Dong, Henghui Ding, Chao Zhang, Hui Qian

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Code can be found at https://github.com/FerdinandZJU/IAP

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21817 2025-07-04 cs.CV 57%

Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping

Weili Zeng, Ziyuan Huang, Kaixiang Ji, Yichao Yan

机构 * MoE Key Lab of Artificial Intelligence, AI Institute Shanghai Jiao Tong University(人工智能大模型关键实验室,上海交通大学AI研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16025 2025-07-04 cs.CV 57%

FeatSharp: Your Vision Model Features, Sharper

Mike Ranzinger, Greg Heinrich, Pavlo Molchanov, Jan Kautz, Bryan Catanzaro, Andrew Tao

机构 * NVIDIA

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments ICML 2025 Version

详情

展开后加载摘要…

URL PDF HTML 收藏