arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46122 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4657 篇

2505.23043 2025-05-30 cs.CV cs.AI 62%

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang, Yu Cheng

机构 * The Chinese University of Hong Kong(香港中文大学) Microsoft(微软公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22200 2025-05-29 cs.CV cs.AI 62%

Investigating Mechanisms for In-Context Vision Language Binding

Darshana Saravanan, Makarand Tapaswi, Vineet Gandhi

机构 * CVIT, IIIT Hyderabad, India(计算机视觉研究所,印度海得拉巴印度理工学院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to MIV at CVPRW 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21547 2025-05-29 cs.CV cs.AI 62%

Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing

Weixing Wang, Zifeng Ding, Jindong Gu, Rui Cao, Christoph Meinel, Gerard de Melo, Haojin Yang

机构 * Hasso Plattner Institute(霍普夫纳研究所) University of Potsdam(波茨坦大学) University of Cambridge(剑桥大学) University of Oxford(牛津大学) German University of Digital Science(德国数字科学大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13062 2025-05-29 cs.MM cs.SD eess.AS 62%

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Yong Ren, Chenxing Li, Le Xu, Hao Gu, Duzhen Zhang, Yujie Chen, Manjie Xu, Ruibo Fu, Shan Yang, Dong Yu

机构 * Chinese Academy of Sciences(中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Tencent AI Lab(腾讯AI实验室) Institute of Automation School of Artificial Intelligence(自动化研究所人工智能学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.MM、eess.AS

Comments Accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20578 2025-05-29 cs.CV cs.AI cs.LG 62%

Interpreting CLIP with Hierarchical Sparse Autoencoders

Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Journal ref Proceedings of the 42st International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19031 2025-05-27 cs.CV cs.AI 62%

Medical Large Vision Language Models with Multi-Image Visual Ability

Xikai Yang, Juzheng Miao, Yuchen Yuan, Jiaze Wang, Qi Dou, Jinpeng Li, Pheng-Ann Heng

机构 * Dept. of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong, China(计算机科学与工程系,香港中文大学,香港,中国) Institute of Medical Intelligence and XR, The Chinese University of Hong Kong, Hong Kong, China(医学智能与XR研究所,香港中文大学,香港,中国)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18434 2025-05-27 cs.CV cs.AI 62%

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP

Yuliang Cai, Jesse Thomason, Mohammad Rostami

机构 * University of Southern California(南加州大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17625 2025-05-26 cs.CL cs.CV 62%

Enhancing Large Vision-Language Models with Layout Modality for Table Question Answering on Japanese Annual Securities Reports

Hayato Aida, Kosuke Takahashi, Takahiro Omi

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted at IIAI AAI 2025, the 3rd International Conference on Computational and Data Sciences in Economics and Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17425 2025-05-26 cs.CV cs.CL 62%

Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads

Wei Jie Yeo, Rui Mao, Moloud Abdar, Erik Cambria, Ranjan Satapathy

机构 * Nanyang Technological University(南洋理工大学) The University of Queensland(昆士兰大学) IHPC(高性能计算中心)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06442 2025-05-21 cs.CV cs.MM 62%

OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection

Yu Liu, Hao Tang, Haiqi Zhang, Jing Qin, Zechao Li

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Centre for Smart Health(智能健康中心)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted to the 34th International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12835 2025-05-20 cs.CL cs.CV 62%

FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models

Hengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li, Zhifeng Gao, Haidong Wang, Zicheng Su, Agachai Sumalee, Renxin Zhong

机构 * School of Intelligent Systems Engineering, Sun Yat-Sen University(中山大学智能系统工程学院) DP Technology(DP技术公司) Beijing University Of Posts and Telecommunications(北京邮电大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Tongji University(同济大学) School of Integrated Innovation, Chulalongkorn University(朱拉隆功大学创新学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07830 2025-05-20 cs.CV cs.AI cs.LG 62%

Captured by Captions: On Memorization and its Mitigation in CLIP Models

Wenhao Wang, Adam Dziedzic, Grace C. Kim, Michael Backes, Franziska Boenisch

机构 * CISPA Georgia Institute of Technology(佐治亚理工学院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11060 2025-05-19 cs.CV cs.AI 62%

CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs

David Méndez, Gianpaolo Bontempo, Elisa Ficarra, Roberto Confalonieri, Natalia Díaz-Rodríguez

机构 * Dept. of Computer Science and Artificial Intelligence, DaSCI Institute, University of Granada(计算机科学与人工智能系,DaSCI研究所,格拉纳达大学) Dept. of Engineering ”Enzo Ferrari”, University of Modena and Reggio Emilia(工程系,摩德纳和雷吉奥艾米利亚大学) Dept. of Mathematics ’Tullio Levi-Civita’, University of Padova(数学系,帕多瓦大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 3 figures, 5 tables. Accepted at IJCNN 2025; to appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10664 2025-05-19 cs.CV cs.AI 62%

CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier

Ziyang Ou

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) University of Rochester(罗切斯特大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figures, not submitted to any conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10453 2025-05-16 cs.CV cs.AI 62%

Vision language models have difficulty recognizing virtual objects

Tyler Tran, Sangeet Khemlani, J. G. Trafton

机构 * US Naval Research Laboratory(美国海军研究实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08910 2025-05-16 cs.CV cs.CL 62%

Behind Maya: Building a Multilingual Vision Language Model

Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, S M Iftekhar Uddin, Shayekh Bin Islam, Roshan Santhosh, Snegha A, Drishti Sharma, Chen Liu, Isha Chaturvedi, Genta Indra Winata, Ashvanth. S, Snehanshu Mukherjee, Alham Fikri Aji

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at VLMs4ALL CVPR 2025 Workshop; corrected workshop name spelling

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09435 2025-05-15 cs.CV cs.AI 62%

Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records

Yili He, Yan Zhu, Peiyao Fu, Ruijie Yang, Tianyi Chen, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang

机构 * Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, China(上海复旦大学基础医学学院数字医学研究中心) University College London, London, UK(伦敦大学学院) Shanghai Key Laboratory of MICCAI, Shanghai, China(上海MICCAI重点实验室) Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai, China(复旦大学中山医院内窥镜中心和内窥镜研究所) Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China(上海内窥镜协同创新中心) Shanghai Institute for Advanced Study of Zhejiang University, Shanghai, China(浙江大学上海高级研究院) Alliance Manchester Business School, The University of Manchester, Manchester, UK(曼彻斯特大学曼彻斯特商业学校) Data Science Institute, Imperial College London, London, UK(伦敦帝国理工学院数据科学研究院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments Early accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09425 2025-05-14 cs.CV cs.CL 62%

Vision-Language Models Do Not Understand Negation

Kumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li, Philip Torr, Yoon Kim, Marzyeh Ghassemi

机构 * Institution1(机构1) Institution2(机构2)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2025; project page: https://negbench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05318 2025-05-09 cs.CV cs.AI cs.CY cs.HC cs.RO 62%

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects

Agnese Chiatti, Sara Bernardini, Lara Shibelski Godoy Piccolo, Viola Schiaffonati, Matteo Matteucci

机构 * Politecnico di Milano, Italy(米兰理工学院,意大利) University of Oxford, UK(牛津大学,英国) CODE University of Applied Sciences, Germany(德国应用科学大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03981 2025-05-09 cs.AI cs.CL cs.LG 62%

X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains

Qianchu Liu, Sheng Zhang, Guanghui Qin, Timothy Ossowski, Yu Gu, Ying Jin, Sid Kiblawi, Sam Preston, Mu Wei, Paul Vozila, Tristan Naumann, Hoifung Poon

机构 * Microsoft Research(微软研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03380 2025-05-07 cs.CV cs.AI eess.IV 62%

Reinforced Correlation Between Vision and Language for Precise Medical AI Assistant

Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang, Marawan Elbatel, Yi Qin, Huijun Hu, Baoxun Li, Wenhui Deng, Weifeng Qin, Hongrui Li, Jialin Liang, Jun Shen, Xiaomeng Li

机构 * Department of Electronic and Computer Engineering, HKUST(香港科技大学电子与计算机工程系) Department of Radiology, Guangdong Provincial Key Laboratory of Malignant Tumor Epigenetics and Gene Regulation, Sun Yat-Sen Memorial Hospital, Sun Yat-Sen University(中山大学放射科、广东省恶性肿瘤表观遗传与基因调控重点实验室、中山纪念医院) Department of Computer Science and Engineering, HKUST(香港科技大学计算机科学与工程系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01958 2025-05-06 cs.CV cs.CL 62%

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models

Liqiang Jing, Guiming Hardy Chen, Ehsan Aghazadeh, Xin Eric Wang, Xinya Du

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Massachusetts at Amherst(马萨诸塞大学阿姆赫斯特分校) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16723 2025-04-24 cs.CV cs.AI 62%

Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering

Ali Anaissi, Junaid Akram, Kunal Chaturvedi, Ali Braytee

机构 * The University of Sydney, School of Computer Science(悉尼大学计算机科学学院) University of Technology Sydney, School of Computer Science(新南威尔士大学技术学院) University of Technology Sydney, TD School(新南威尔士大学TD学院) Australian Catholic University, Peter Faber Business School(澳大利亚天主教大学彼得·法伯商学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages, 2 figures, 2025 International Conference on Computational Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15199 2025-04-22 cs.CV cs.AI cs.LG cs.PF 62%

Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning

Yassir Benhammou, Alessandro Tiberio, Gabriel Trautmann, Suman Kalyan

机构 * NstarX

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 2 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14848 2025-04-22 cs.CV cs.AI 62%

Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation

Yunpu Zhao, Rui Zhang, Junbin Xiao, Ruibo Hou, Jiaming Guo, Zihao Zhang, Yifan Hao, Yunji Chen

机构 * University of Science and Technology of China(中国科学技术大学) SKL of Processors, Institute of Computing Technology, CAS(中国科学院计算技术研究所处理器专项实验室) National University of Singapore(新加坡国立大学) University of Chinese Academy of Sciences(中国科学院大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12256 2025-04-17 cs.CV cs.AI cs.LG 62%

FLIP Reasoning Challenge

Andreas Plesner, Turlan Kuzhagaliyev, Roger Wattenhofer

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Published at First Workshop on Open Science for Foundation Models at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11038 2025-04-16 cs.CV cs.AI 62%

QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models

Yudong Zhang, Ruobing Xie, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Yu Wang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by NAACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20826 2025-04-16 cs.CV cs.CL cs.LG eess.IV 62%

Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Yucong Meng, Kexue Fu, Feilong Tang, Shuo Wang, Zhijian Song

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09480 2025-04-15 cs.CV cs.AI 62%

Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation

Yongchao Feng, Yajie Liu, Shuai Yang, Wenrui Cai, Jinqing Zhang, Qiqi Zhan, Ziyue Huang, Hongxi Yan, Qiao Wan, Chenguang Liu, Junzhe Wang, Jiahui Lv, Ziqi Liu, Tengyuan Shi, Qingjie Liu, Yunhong Wang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments A Review and Evaluation about Vision-Language Model for Object Detection and Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09203 2025-04-15 cs.CV cs.AI 62%

AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images

Saikat Dutta, Akhil Vasim, Siddhant Gole, Hamid Rezatofighi, Biplab Banerjee

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted at EarthVision workshop, CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏