arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4657 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4657 篇

2506.04743 2025-11-18 cs.CV 57%

SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMs

Shuhan Xu, Siyuan Liang, Hongling Zheng, Aishan Liu, Xinbiao Wang, Yong Luo, Fu Lin, Leszek Rutkowski, Dacheng Tao

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09443 2025-11-18 cs.CL cs.LG 57%

Scaling Laws for Conditional Emergence of Multilingual Image Captioning via Generalization from Translation

Julian Spravil, Sebastian Houben, Sven Behnke

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11704 2025-11-18 cs.LG cs.CV 57%

Simple Vision-Language Math Reasoning via Rendered Text

Matvey Skripkin, Elizaveta Goncharova, Andrey Kuznetsov

机构 * FusionBrain Lab(融合脑实验室) HSE University(俄罗斯高等经济学院) Innopolis University(因诺波利斯大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08283 2025-11-18 cs.LG cs.CV 57%

DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities

Jueqing Lu, Yuanyuan Qi, Xiaohao Yang, Shuaicheng Niu, Fucai Ke, Shujie Zhou, Wei Tan, Jionghao Lin, Wray Buntine, Hamid Rezatofighi, Lan Du

机构 * Monash University(墨尔本大学) Nanyang Technological University(南洋理工大学) The University of Hong Kong(香港大学) VinUni

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Updates to v1. Added new coauthors and extended the experimental section

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11141 2025-11-17 cs.CL cs.CY cs.LG 57%

PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases

Udo Schlegel, Franziska Weeber, Jian Lan, Thomas Seidl

机构 * Ludwig-Maximilians-University Munich(慕尼黑路德维希-马克西米利安大学) Munich Center for Machine Learning(慕尼黑机器学习中心) University of Stuttgart(斯图加特大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments 8 pages, accpeted as short paper at MMM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10769 2025-11-17 cs.CV 57%

Unifying Segment Anything in Microscopy with Vision-Language Knowledge

Manyu Li, Ruian He, Zixian Zhang, Chenxi Ma, Weimin Tan, Bo Yan

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10774 2025-11-17 cs.CV 57%

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

Junjie Zhang, Feng Zhao, Hanqiang Liu, Jun Yu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08172 2025-11-17 cs.AI 57%

An Efficient Training Pipeline for Reasoning Graphical User Interface Agents

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10098 2025-11-14 cs.CV 57%

MTAttack: Multi-Target Backdoor Attacks against Large Vision-Language Models

Zihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng, Xiao Bai

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments AAAI2026, with supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09883 2025-11-14 cs.CV 57%

HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models

Liheng Zhang, Jin Wang, Hui Li, Bingfeng Zhang, Weifeng Liu

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09868 2025-11-14 cs.CV 57%

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

Peng Gao, Yujian Lee, Xiaofeng Zhang, Zailong Chen, Hui Zhang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09064 2025-11-13 cs.CV 57%

Diversifying Counterattacks: Orthogonal Exploration for Robust CLIP Inference

Chengze Jiang, Minjing Dong, Xinli Shi, Jie Gui

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to AAAI-2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08909 2025-11-13 cs.CV 57%

Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images

Zimao Lu, Hui Xu, Bing Liu, Ke Wang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07068 2025-11-11 cs.CV cs.LG 57%

ClusterMine: Robust Label-Free Visual Out-Of-Distribution Detection via Concept Mining from Text Corpora

Nikolas Adaloglou, Diana Petrusheva, Mohamed Asker, Felix Michels, Markus Kollmann

机构 * Heinrich Heine University of Düsseldorf(海因里希-海涅大学杜塞尔多夫分校)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted in WACV 2026. Code in https://github.com/HHU-MMBS/clustermine_wacv_official 9 Tables, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05642 2025-11-11 cs.RO cs.AR cs.CV cs.SY eess.SY 57%

Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots

Justin Williams, Kishor Datta Gupta, Roy George, Mrinmoy Sarkar

机构 * Department of Cyber-Physical Systems, Clark Atlanta University(克雷克阿特拉大学计算机物理系统系) Siemens Corporation(西门子公司)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08668 2025-11-06 cs.CV 57%

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Songtao Jiang, Yuan Wang, Sibo Song, Tianxiang Hu, Chenyi Zhou, Bin Pu, Yan Zhang, Zhibo Yang, Yang Feng, Joey Tianyi Zhou, Jin Hao, Zijian Chen, Ruijia Wu, Tao Tang, Junhui Lv, Hongxia Xu, Hongwei Wang, Jun Xiao, Bin Feng, Fudong Zhu, Kenli Li, Weidi Xie, Jimeng Sun, Jian Wu, Zuozhu Liu

机构 * College of Computer Science and Technology, Zhejiang University-University of Illinois Urbana-Champaign Institute(浙江大学计算机科学与技术学院) Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine(浙江大学口腔医院) Alibaba Inc(阿里巴巴集团) College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) Angelalign Technology Inc.(Angelalign技术有限公司) CFAR & IHPC, Agency for Science, Technology and Research(CFAR与IHPC,新加坡科技研究局) Department of Orthodontics, Shanghai Ninth People’s Hospital, College of Stomatology, Shanghai Jiao Tong University(上海第九人民医院正畸科,上海交通大学口腔医学院)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11772 2025-11-06 cs.CV cs.LG 57%

CLIP Meets Diffusion: A Synergistic Approach to Anomaly Detection

Byeongchan Lee, John Won, Seunghyun Lee, Jinwoo Shin

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国先进科学技术研究院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted at TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25067 2025-11-05 cs.CV 57%

DRIP: Dynamic patch Reduction via Interpretable Pooling

Yusen Peng, Sachin Kumar

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Need more refinement

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01694 2025-11-04 cs.LG cs.AI 57%

Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering

Hossein Abdi, Mingfei Sun, Wei Pan

机构 * The University of Manchester(曼彻斯特大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01550 2025-11-04 cs.AI 57%

Analyzing Sustainability Messaging in Large-Scale Corporate Social Media

Ujjwal Sharma, Stevan Rudinac, Ana Mićković, Willemijn van Dolen, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07409 2025-11-04 cs.CV cs.LG 57%

MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification

Anh-Tien Nguyen, Duy Minh Ho Nguyen, Nghiem Tuong Diep, Trung Quoc Nguyen, Nhat Ho, Jacqueline Michelle Metsch, Miriam Cindy Maurer, Daniel Sonntag, Hanibal Bohnenberger, Anne-Christin Hauschild

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Published in Transactions on Machine Learning Research (09/2025)

Journal ref Transactions on Machine Learning Research (09/2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00446 2025-11-04 cs.CV cs.CR cs.LG 57%

ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training

Xin Yao, Haiyang Zhao, Yimin Chen, Jiawei Guo, Kecheng Huang, Ming Zhao

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) Miner School of Computer & Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校Miner计算机与信息科学学院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02866 2025-11-04 cs.CV 57%

OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery

Xiucheng Liang, Jinheng Xie, Tianhong Zhao, Rudi Stouffs, Filip Biljecki

机构 * Department of Architecture, National University of Singapore(建筑系,新加坡国立大学) Department of Electrical and Computer Engineering, National University of Singapore(电气与计算机工程系,新加坡国立大学) School of Artificial Intelligence, Shenzhen Technology University(人工智能学院,深圳科技大学) Department of Real Estate, National University of Singapore(房地产系,新加坡国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Journal ref ISPRS Journal of Photogrammetry and Remote Sensing 230: 918-942, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08478 2025-11-04 cs.IR cs.AI cs.LG 57%

Federated Vision-Language-Recommendation with Personalized Fusion

Zhiwei Li, Guodong Long, Jing Jiang, Chengqi Zhang, Qiang Yang

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

Comments 15 pages, 10 figures, 7 tables, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25179 2025-10-30 cs.AI 57%

Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models

Juan Ren, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University, Australia(计算机学院,麦考瑞大学,澳大利亚)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25175 2025-10-30 cs.CV 57%

Test-Time Adaptive Object Detection with Foundation Model

Yingjie Gao, Yanan Zhang, Zhi Cai, Di Huang

机构 * State Key Laboratory of Complex and Critical Software Environment, Beihang University(复杂与关键软件环境国家重点实验室,北京航空航天大学) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25051 2025-10-30 cs.CV cs.LG 57%

Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models

Shunjie-Fabian Zheng, Hyeonjun Lee, Thijs Kooi, Ali Diba

机构 * Department of Medicine I, LMU University Hospital, LMU Munich, Germany(慕尼黑大学医学院第一医学部,LMU慕尼黑大学医院) Lunit Inc.(Lunit公司)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to Computer Vision for Automated Medical Diagnosis (CVAMD) Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24650 2025-10-29 cs.AI 57%

Advancing site-specific disease and pest management in precision agriculture: From reasoning-driven foundation models to adaptive, feedback-based learning

Nitin Rai, Daeun, Choi, Nathan S. Boyd, Arnold W. Schumann

机构 * Department of Horticultural Sciences(园艺科学系) Gulf Coast Research and Education Center(墨西哥湾沿岸研究与教育中心) University of Florida(佛罗里达大学) Department of Agricultural and Biological Engineering(农业与生物工程系) Department of Soil, Water, and Ecosystem Sciences(土壤、水与生态系统科学系) Citrus Research and Education Center(柑橘研究与教育中心)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.AI

Comments 26 pages, 8 figures, and 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07046 2025-10-29 cs.CV 57%

RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning

Yunchuan Ma, Laiyun Qing, Guorong Li, Yuankai Qi, Amin Beheshti, Quan Z. Sheng, Qingming Huang

机构 * University of Chinese Academy of Science, Beijing,100190, China(中国科学院大学) Macquarie University(麦考瑞大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Published in Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23894 2025-10-29 cs.CV 57%

Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation

Jinxin Zhou, Jiachen Jiang, Zhihui Zhu

机构 * Department of Computer Science(计算机科学系)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 23 pages, 10 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏