arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2507.11055 2025-07-22 cs.CV 57%

Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic Approximation

Shuchang Ye, Usman Naseem, Mingyuan Meng, Jinman Kim

机构 * The University of Sydney(悉尼大学) Macquarie University(麦觉瑞大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09222 2025-07-22 cs.CV cs.LG 57%

Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift

Behraj Khan, Tahir Qasim Syed, Nouman M. Durrani, Bilal Naseem, Shabir Ahmad, Rizwan Qureshi

机构 * Institute of Business Administration Karachi(Karachi商业管理学院) National University of Computer and Emerging Sciences(国家计算机与新兴科学大学) CAIMI Pvt Ltd(CAIMI私营有限公司) Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,佛罗里达大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10283 2025-07-22 cs.RO cs.AI cs.SY eess.IV eess.SY 57%

ASMA: An Adaptive Safety Margin Algorithm for Vision-Language Drone Navigation via Scene-Aware Control Barrier Functions

Sourav Sanyal, Kaushik Roy

机构 * School of Electrical and Computer Engineering, Purdue University(电气与计算机工程学院,普渡大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11200 2025-07-21 cs.CV 57%

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

Che Liu, Jiazhen Pan, Weixiang Shen, Wenjia Bai, Daniel Rueckert, Rossella Arcucci

机构 * Imperial College London, UK(伦敦帝国学院) Technical University of Munich, Germany(慕尼黑技术大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13359 2025-07-21 cs.CV 57%

Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives

Yang Zhou, Junjie Li, CongYang Ou, Dawei Yan, Haokui Zhang, Xizhe Xue

机构 * Northwestern Polytechnical University(北华理工大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments 27 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13113 2025-07-18 cs.CV 57%

Leveraging Language Prior for Infrared Small Target Detection

Pranav Singh, Pravendra Singh

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Roorkee(计算机科学与工程系,印度理工学院Roorkee)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12755 2025-07-18 cs.CV cs.LG 57%

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation

Yanchen Guan, Haicheng Liao, Chengyue Wang, Bonan Wang, Jiaxun Zhang, Jia Hu, Zhenning Li

机构 * State Key Laboratory of Internet of Things for Smart City(物联网智能城市国家重点实验室) University of Macau(澳门大学) Department of Civil Engineering(土木工程系) Department of Computer and Information Science(计算机与信息科学系) College of Transportation Engineering(交通工程学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03770 2025-07-18 cs.CR cs.AI 57%

JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model

Yi Nian, Shenzhe Zhu, Yuehan Qin, Li Li, Ziyi Wang, Chaowei Xiao, Yue Zhao

机构 * University of Southern California(南加州大学) University of Toronto(多伦多大学) University of Maryland(马里兰大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19697 2025-07-18 cs.CV 57%

Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion

Yuan Bian, Min Liu, Yunqi Yi, Xueping Wang, Yaonan Wang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) National Engineering Research Center of Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) College of Information Science and Engineering, Hunan Normal University(湖南师范大学信息科学与工程学院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11102 2025-07-16 cs.CV 57%

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

Jie Yang, Wang Zeng, Sheng Jin, Lumin Xu, Wentao Liu, Chen Qian, Zhen Li, Ruimao Zhang

机构 * Sun Yat-sen University, Shenzhen(中山大学深圳校区) Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区) SenseTime Research(商汤科技研究院) Chinese University of Hong Kong(香港中文大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Extended Version of KptLLM. arXiv admin note: text overlap with arXiv:2411.01846

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05464 2025-07-16 cs.CL 57%

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Shiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu, Siyang Gao, Miao Xiong, Manling Li, Junxian He

机构 * City University of Hong Kong(香港城市大学) Hong Kong University of Science(香港科学大学) National University of Singapore(新加坡国立大学) Northwestern University(西北大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments ICML 2025. Camera-ready version updated. Our code is publicly available at https://github.com/shiqichen17/VLM_Merging

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01069 2025-07-15 cs.CV 57%

UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment

Hantao Zhou, Longxiang Tang, Rui Yang, Guanyi Qin, Yan Zhang, Yutao Li, Xiu Li, Runze Hu, Guangtao Zhai

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Media Analytics and Computing Lab, Department of Artificial Intelligence, School of Informatics, Xiamen University(媒体分析与计算实验室,人工智能系,厦门大学) School of Computer Science and Technology, Ocean University of China(计算机科学与技术学院,中国海洋大学) Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(图像通信与信息处理研究所,上海交通大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09615 2025-07-15 cs.CV 57%

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score

Eman Ali, Sathira Silva, Chetan Arora, Muhammad Haris Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫莫德·本·扎耶德人工智能大学) IIT Delhi(德里印度理工学院) Alexandria University(亚历山大大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09008 2025-07-15 cs.CV 57%

VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels

Xiwei Xuan, Xiaoqi Wang, Wenbin He, Jorge Piazentin Ono, Liang Gou, Kwan-Liu Ma, Liu Ren

机构 * Department of Computer Science, University of California, Davis, CA, USA(加州大学戴维斯分校计算机科学系) Bosch Center for Artificial Intelligence (BCAI), Bosch Research North America(博世人工智能中心(BCAI)、博世北美研究部) Splunk Technology, San Jose, CA, USA(Splunk技术公司)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments IEEE Transactions on Visualization and Computer Graphics (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08982 2025-07-15 eess.IV cs.CV cs.LG 57%

VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models

Hanene F. Z. Brachemi Meftah, Wassim Hamidouche, Sid Ahmed Fezza, Olivier Déforges

机构 * Univ. Rennes, INSA Rennes, CNRS, IETR - UMR 6164(里昂大学、里昂国家理工学院、国家科学研究中心、IETR - UMR 6164)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07104 2025-07-14 cs.CV 57%

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models

Tiezheng Zhang, Yitong Li, Yu-cheng Chou, Jieneng Chen, Alan Yuille, Chen Wei, Junfei Xiao

机构 * Johns Hopkins University(约翰霍普金斯大学) Tsinghua University(清华大学) Rice University(Rice 大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Project Page: https://lambert-x.github.io/Vision-Language-Vision/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07562 2025-07-11 cs.CL 57%

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs

Jierun Chen, Tiezheng Yu, Haoli Bai, Lewei Yao, Jiannan Wu, Kaican Li, Fei Mi, Chaofan Tao, Lei Zhu, Manyi Zhang, Xiaohui Li, Lu Hou, Lifeng Shang, Qun Liu

机构 * Huawei Technologies(华为技术有限公司) HKUST(香港科技大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07148 2025-07-11 cs.CV cs.LG 57%

Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey

Getamesay Haile Dagnaw, Yanming Zhu, Muhammad Hassan Maqsood, Wencheng Yang, Xingshuai Dong, Xuefei Yin, Alan Wee-Chung Liew

机构 * Griffith University(格里菲斯大学) University of Southern Queensland(南部昆士兰大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17670 2025-07-09 cs.LG cs.AI 57%

Towards General Continuous Memory for Vision-Language Models

Wenyi Wu, Zixuan Song, Kun Zhou, Yifei Shao, Zhiting Hu, Biwei Huang

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04699 2025-07-08 cs.CV 57%

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

Zexi Jia, Chuanwei Huang, Hongyan Fei, Yeshuang Zhu, Zhiqiang Yuan, Ying Deng, Jiapei Zhang, Jinchao Zhang, Jie Zhou

机构 * WeChat AI, Tencent Inc, China(腾讯公司)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05163 2025-07-08 cs.CV cs.LG 57%

Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models

Aishwarya Venkataramanan, Paul Bodesheim, Joachim Denzler

机构 * Computer Vision Group, Friedrich Schiller University Jena(计算机视觉组,费迪里奇·施勒尔大学耶纳)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments UAI 2025, 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20612 2025-07-08 cs.CV 57%

IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware Prompting

Hao Fu, Hanbin Zhao, Jiahua Dong, Henghui Ding, Chao Zhang, Hui Qian

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Code can be found at https://github.com/FerdinandZJU/IAP

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21817 2025-07-04 cs.CV 57%

Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping

Weili Zeng, Ziyuan Huang, Kaixiang Ji, Yichao Yan

机构 * MoE Key Lab of Artificial Intelligence, AI Institute Shanghai Jiao Tong University(人工智能大模型关键实验室,上海交通大学AI研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16025 2025-07-04 cs.CV 57%

FeatSharp: Your Vision Model Features, Sharper

Mike Ranzinger, Greg Heinrich, Pavlo Molchanov, Jan Kautz, Bryan Catanzaro, Andrew Tao

机构 * NVIDIA

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments ICML 2025 Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03015 2025-07-03 cs.RO cs.CV 57%

Balancing Performance and Efficiency in Zero-shot Robotic Navigation

Dmytro Kuzmenko, Nadiya Shvai

机构 * Department of Multimedia Systems, National University of Kyiv-Mohyla Academy, Kyiv, Ukraine(多媒体系统系,基辅-莫希拉学院国家大学,乌克兰基辅) Department of Mathematics, National University of Kyiv-Mohyla Academy, Kyiv, Ukraine(数学系,基辅-莫希拉学院国家大学,乌克兰基辅)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Submitted to ICTERI 2024 Posters Track

Journal ref ICTERI 2024: Communications in Computer and Information Science, vol. 2020, pp. 370-381, Springer, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00886 2025-07-02 cs.CV cs.RO 57%

GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond

Anna-Maria Halacheva, Jan-Nico Zaech, Xi Wang, Danda Pani Paudel, Luc Van Gool

机构 * ETH Zurich(苏黎世联邦理工学院) TU Munich(慕尼黑技术大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24102 2025-07-01 cs.CV 57%

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Xiangtai Li, Tao Zhang, Yanwei Li, Haobo Yuan, Shihao Chen, Yikang Zhou, Jiahao Meng, Yueyi Sun, Shilin Xu, Lu Qi, Tianheng Cheng, Yi Lin, Zilong Huang, Wenhao Huang, Jiashi Feng, Guang Shi

机构 * ByteDance Seed(字节跳动种子) Wuhan University(武汉大学) Peking University(北京大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Datasets and Models: https://github.com/lxtGH/DenseWorld-1M

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22979 2025-07-01 cs.CV 57%

Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation

Jie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke, Efstratios Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) Singapore Management University(新加坡管理大学) The Netherlands Cancer Institute(荷兰癌症研究院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments ICCV2025 Proceeding

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07503 2025-07-01 cs.CV 57%

Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts

Shiu-hong Kao, Yu-Wing Tai, Chi-Keung Tang

机构 * HKUST(香港理工大学) Dartmouth College(达特茅斯学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Project page: https://danielshkao.github.io/thinkfirst.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22395 2025-06-30 cs.CV 57%

Test-Time Consistency in Vision Language Models

Shih-Han Chou, Shivam Chandhok, James J. Little, Leonid Sigal

机构 * Department of Computer Science, University of British Columbia(不列颠哥伦比亚大学计算机科学系) Vector Institute for AI(人工智能矢量研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏