arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2504.07934 2025-06-02 cs.CV 57%

SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Xiyao Wang, Zhengyuan Yang, Chao Feng, Hongjin Lu, Linjie Li, Chung-Ching Lin, Kevin Lin, Furong Huang, Lijuan Wang

机构 * University of Maryland College Park(马里兰大学学院公园分校) Microsoft(微软) University of Michigan(密歇根大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 27 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23766 2025-05-30 cs.CV 57%

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

Yunze Man, De-An Huang, Guilin Liu, Shiwei Sheng, Shilong Liu, Liang-Yan Gui, Jan Kautz, Yu-Xiong Wang, Zhiding Yu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) NVIDIA

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2025. Project Page: https://yunzeman.github.io/argus/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23040 2025-05-30 cs.CV 57%

Deep Modeling and Optimization of Medical Image Classification

Yihang Wu, Muhammad Owais, Reem Kateb, Ahmad Chaddad

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in ISBI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18686 2025-05-30 cs.CV 57%

WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation

Yang Liu, Silin Cheng, Xinwei He, Sebastien Ourselin, Lei Tan, Gen Luo

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24129 2025-05-30 cs.CV cs.LG 57%

It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data

Dominik Schnaus, Nikita Araslanov, Daniel Cremers

机构 * TU Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted to CVPR 2025, Project page: https://dominik-schnaus.github.io/itsamatch/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22209 2025-05-29 cs.CV 57%

A Survey on Training-free Open-Vocabulary Semantic Segmentation

Naomi Kombol, Ivan Martinović, Siniša Šegvić

机构 * Sveučilište u Zagrebu(扎格雷布大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22079 2025-05-29 cs.CV 57%

Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis

Hanbin Ko, Chang-Min Park

机构 * Interdisciplinary Program in Bioengineering, Seoul National University Graduate School(生物工程跨学科项目,首尔国立大学研究生院) Integrated Major in Innovative Medical Science, Seoul National University Graduate School(创新医学联合专业,首尔国立大学研究生院) Department of Radiology, Seoul National University Hospital(放射科,首尔国立大学医院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 16 pages (8 main, 2 references, 6 appendix), 13 figures. Accepted to CVPR 2025. This author-accepted manuscript includes an expanded ethics/data user agreement section. The final version will appear in the Proceedings of CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21967 2025-05-29 cs.CL 57%

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20362 2025-05-28 cs.IR cs.AI 57%

VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration

Jiahui Geng, Qing Li, Zongxiong Chen, Yuxia Wang, Derui Zhu, Zhuohan Xie, Chenyang Lyu, Xiuying Chen, Preslav Nakov, Fakhri Karray

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Fraunhofer Institute for Open Communication Systems(弗劳恩霍夫开放通信系统研究所) Technical University of Munich(慕尼黑技术大学) Alibaba International Digital Commerce(阿里巴巴国际数字商务)

专题命中 图文多模态 :image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12635 2025-05-28 cs.CV 57%

Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning

Yunhao Gou, Hansi Yang, Zhili Liu, Kai Chen, Yihan Zeng, Lanqing Hong, Zhenguo Li, Qun Liu, Bo Han, James T. Kwok, Yu Zhang

机构 * Southern University of Science and Technology(南方科技大学) The Hong Kong University of Science and Technology(香港科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20236 2025-05-27 cs.CV 57%

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

Weihao Xuan, Qingcheng Zeng, Heli Qi, Junjue Wang, Naoto Yokoya

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19373 2025-05-27 cs.CV 57%

DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models

Niloufar Alipour Talemi, Hossein Kashiani, Hossein R. Nowdeh, Fatemeh Afghah

机构 * Clemson University(克莱姆森大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted at the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19242 2025-05-27 cs.CV 57%

Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model

Alaa Dalaq, Muzammil Behzad

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19139 2025-05-27 cs.CV 57%

The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework

Feiran Liu, Yuzhe Zhang, Xinyi Huang, Yinan Peng, Xinfeng Li, Lixu Wang, Yutong Shen, Ranjie Duan, Simeng Qin, Xiaojun Jia, Qingsong Wen, Wei Dong

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Technology(北京理工大学) Hengxin Tech(恒心科技) Alibaba Group(阿里巴巴集团) Squirrel Ai Learning(squirrel AI 学习)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18719 2025-05-27 cs.RO cs.AI 57%

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Guanxing Lu, Wenkai Guo, Chubin Zhang, Yuheng Zhou, Haonan Jiang, Zifeng Gao, Yansong Tang, Ziwei Wang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18594 2025-05-27 cs.CV cs.IR 57%

EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models

GuangHao Meng, Sunan He, Jinpeng Wang, Tao Dai, Letian Zhang, Jieming Zhu, Qing Li, Gang Wang, Rui Zhang, Yong Jiang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00965 2025-05-27 cs.CV cs.LG 57%

CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling

Xinze Wang, Chen Chen, Yinfei Yang, Hong-You Chen, Bowen Zhang, Aditya Pal, Xiangxin Zhu, Xianzhi Du

机构 * Apple(苹果公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18039 2025-05-26 cs.CV 57%

Clip4Retrofit: Enabling Real-Time Image Labeling on Edge Devices via Cross-Architecture CLIP Distillation

Li Zhong, Ahmed Ghazal, Jun-Jun Wan, Frederik Zilly, Patrick Mackens, Joachim E. Vollrath, Bogdan Sorin Coseriu

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17835 2025-05-26 cs.CV 57%

VLM Models and Automated Grading of Atopic Dermatitis

Marc Lalonde, Hamed Ghodrati

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14999 2025-05-23 cs.CV 57%

Leveraging Habitat Information for Fine-grained Bird Identification

Tin Nguyen, Peijie Chen, Anh Totti Nguyen

机构 * Auburn University(阿伯拉罕大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13426 2025-05-20 cs.CV 57%

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

Liang Chen, Hongcheng Gao, Tianyu Liu, Zhiqi Huang, Flood Sung, Xinyu Zhou, Yuxin Wu, Baobao Chang

机构 * Peking University(北京大学) UCAS(中国科学院大学) Moonshot AI

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 21 pages, 14 figures, code released at https://github.com/chenllliang/G1

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12670 2025-05-20 cs.CV 57%

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning

Lihong Chen, Hossein Hassani, Soodeh Nikan

机构 * Department of Electrical and Computer Engineering, Western University(电气与计算机工程系,西方大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09120 2025-05-20 cs.CL 57%

Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues

Ye-eun Cho, Yunho Maeng

机构 * Sungkyunkwan University(顺天妇女大学) Ewha Womans University & LLM Experimental Lab, MODULABS(成均馆大学及LLM实验实验室,MODULABS)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments 11 pages, 4 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11121 2025-05-19 cs.CV 57%

Redundancy-Aware Pretraining of Vision-Language Foundation Models in Remote Sensing

Mathis Jürgen Adler, Leonard Hackel, Gencer Sumbul, Begüm Demir

机构 * TU Berlin(柏林技术大学) BIFOLD(BIFOLD机构) EPFL(苏黎世联邦理工学院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted at IEEE International Geoscience and Remote Sensing Symposium (IGARSS) 2025. Our code is available at https://git.tu-berlin.de/rsim/redundacy-aware-rs-vlm

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10533 2025-05-16 cs.CV cs.LG 57%

Enhancing Multi-Image Question Answering via Submodular Subset Selection

Aaryan Sharma, Shivansh Gupta, Samar Agarwal, Vishak Prasad C., Ganesh Ramakrishnan

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10289 2025-05-16 cs.CV 57%

MSCI: Addressing CLIP's Inherent Limitations for Compositional Zero-Shot Learning

Yue Wang, Shuai Xu, Xuelin Zhu, Yicong Li

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Key Laboratory of Social Computing and Cognitive Intelligence (Dalian University of Technology), Ministry of Education, China(社会科学计算与认知智能重点实验室(大连理工大学)) The Hong Kong Polytechnic University(香港理工大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12514 2025-05-14 cs.RO cs.CV 57%

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Kun Wu, Zhiyuan Xu, Ning Liu, Ran Cheng, Chaomin Shen, Yaxin Peng, Feifei Feng, Jian Tang

机构 * East China Normal University(东华师范大学) Midea Group, AI Lab(美的集团人工智能实验室) Syracuse University(雪城大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) Shanghai University(上海大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments add more citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06840 2025-05-13 cs.CV 57%

Visual Instruction Tuning with Chain of Region-of-Interest

Yixin Chen, Shuai Zhang, Boran Han, Bernie Wang

机构 * Amazon Web Services(亚马逊网络服务)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments N/A

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11370 2025-05-13 cs.CV 57%

Transmission Line Defect Detection Based on UAV Patrol Images and Vision-language Pretraining

Ke Zhang, Zhaoye Zheng, Yurong Guo, Jiacun Wang, Jiyuan Yang, Yangjie Xiao

机构 * Department of Electronic and Communication Engineering, North China Electric Power University(电子与通信工程系,华北电力大学) Hebei Key Laboratory of Power Internet of Things Technology, North China Electric Power University(河北省电力物联网技术重点实验室,华北电力大学) Computer Science and Software Engineering Department, Monmouth University(计算机科学与软件工程系,蒙特莫恩大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08837 2025-05-09 cs.LG cs.AI 57%

VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Haozhe Wang, Chao Qu, Zuming Huang, Wei Chu, Fangzhen Lin, Wenhu Chen

机构 * HKUST(香港科技大学) University of Waterloo(滑铁卢大学) Vector Institute(向量研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏