arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-13 至 2025-08-13 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2508.07414 2025-08-13 cs.CL cs.LG 83%

Grounding Multilingual Multimodal LLMs With Cultural Knowledge

Jean de Dieu Nyandwi, Yueqi Song, Simran Khanuja, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09087 2025-08-13 cs.CV 70%

Addressing Bias in VLMs for Glaucoma Detection Without Protected Attribute Supervision

Ahsan Habib Akash, Greg Murray, Annahita Amireskandari, Joel Palko, Carol Laxson, Binod Bhattarai, Prashnna Gyawali

机构 * West Virginia University(西弗吉尼亚大学) University of Aberdeen(阿伯丁大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments 3rd Workshop in Data Engineering in Medical Imaging (DEMI), MICCAI-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08939 2025-08-13 cs.CV 70%

MADPromptS: Unlocking Zero-Shot Morphing Attack Detection with Multiple Prompt Aggregation

Eduarda Caldeira, Fadi Boutros, Naser Damer

机构 * Fraunhofer IGD and Department of Computer Science, TU Darmstadt(弗劳恩霍夫研究所(IGD)和图宾根大学计算机科学系)

专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted at ACM Multimedia Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08926 2025-08-13 cs.AI 70%

Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models

Wei Cai, Jian Zhao, Yuchu Jiang, Tianle Zhang, Xuelong Li

机构 * Peking University(北京大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Northwestern Polytechnical University(西北工业大学) Southeast University(东南大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08644 2025-08-13 cs.CV 70%

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation

Guiming Cao, Yuming Ou

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02865 2025-08-13 eess.IV cs.AI cs.CL cs.CV 67%

VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge

Zihan Li, Diping Song, Zefeng Yang, Deming Wang, Fei Li, Xiulan Zhang, Paul E. Kinahan, Yu Qiao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Washington(华盛顿大学) Shenzhen Institutes of Advanced Technology(深圳先进技术研究所) Chinese Academy of Sciences(中国科学院) State Key Laboratory of Ophthalmology(眼科学国家重点实验室) Zhongshan Ophthalmic Center(中山眼科中心) Sun Yat-sen University(中山大学) Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science(广东省眼科学与视觉科学重点实验室) Guangdong Provincial Clinical Research Center for Ocular Diseases(广东省眼科临床研究中心) Department of Bioengineering(生物工程系) Department of Radiology(放射科)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by IEEE TPAMI, 14 pages, 15 tables, 4 figures with Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06795 2025-08-13 cs.CL cs.CV 62%

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

Yuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang, Zhengwei Fang, Jingyuan Zhang, Jiawei Chen, Zinan Liu, Yu Tian

机构 * University of Chinese Academy of Sciences(中国科学院大学) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学) Gaoling School of Artificial Intelligence, Renmin University of China(人工智能学院,中国人民大学) Kuaishou Technology Inc.(快手科技有限公司) Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University(多信息处理重点实验室,华东师范大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03926 2025-08-13 cs.CV 57%

Multiple Stochastic Prompt Tuning for Few-shot Adaptation under Extreme Domain Shift

Debarshi Brahma, Soma Biswas

机构 * Indian Institute of Science(印度科学研究院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏