arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-13 至 2025-08-13 共收录 50 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2508.07414 2025-08-13 cs.CL cs.LG 83%

Grounding Multilingual Multimodal LLMs With Cultural Knowledge

Jean de Dieu Nyandwi, Yueqi Song, Simran Khanuja, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09087 2025-08-13 cs.CV 70%

Addressing Bias in VLMs for Glaucoma Detection Without Protected Attribute Supervision

Ahsan Habib Akash, Greg Murray, Annahita Amireskandari, Joel Palko, Carol Laxson, Binod Bhattarai, Prashnna Gyawali

机构 * West Virginia University(西弗吉尼亚大学) University of Aberdeen(阿伯丁大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments 3rd Workshop in Data Engineering in Medical Imaging (DEMI), MICCAI-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08939 2025-08-13 cs.CV 70%

MADPromptS: Unlocking Zero-Shot Morphing Attack Detection with Multiple Prompt Aggregation

Eduarda Caldeira, Fadi Boutros, Naser Damer

机构 * Fraunhofer IGD and Department of Computer Science, TU Darmstadt(弗劳恩霍夫研究所(IGD)和图宾根大学计算机科学系)

专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted at ACM Multimedia Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08926 2025-08-13 cs.AI 70%

Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models

Wei Cai, Jian Zhao, Yuchu Jiang, Tianle Zhang, Xuelong Li

机构 * Peking University(北京大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Northwestern Polytechnical University(西北工业大学) Southeast University(东南大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08644 2025-08-13 cs.CV 70%

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation

Guiming Cao, Yuming Ou

专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02865 2025-08-13 eess.IV cs.AI cs.CL cs.CV 67%

VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge

Zihan Li, Diping Song, Zefeng Yang, Deming Wang, Fei Li, Xiulan Zhang, Paul E. Kinahan, Yu Qiao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Washington(华盛顿大学) Shenzhen Institutes of Advanced Technology(深圳先进技术研究所) Chinese Academy of Sciences(中国科学院) State Key Laboratory of Ophthalmology(眼科学国家重点实验室) Zhongshan Ophthalmic Center(中山眼科中心) Sun Yat-sen University(中山大学) Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science(广东省眼科学与视觉科学重点实验室) Guangdong Provincial Clinical Research Center for Ocular Diseases(广东省眼科临床研究中心) Department of Bioengineering(生物工程系) Department of Radiology(放射科)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by IEEE TPAMI, 14 pages, 15 tables, 4 figures with Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06795 2025-08-13 cs.CL cs.CV 62%

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

Yuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang, Zhengwei Fang, Jingyuan Zhang, Jiawei Chen, Zinan Liu, Yu Tian

机构 * University of Chinese Academy of Sciences(中国科学院大学) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学) Gaoling School of Artificial Intelligence, Renmin University of China(人工智能学院,中国人民大学) Kuaishou Technology Inc.(快手科技有限公司) Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University(多信息处理重点实验室,华东师范大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03926 2025-08-13 cs.CV 57%

Multiple Stochastic Prompt Tuning for Few-shot Adaptation under Extreme Domain Shift

Debarshi Brahma, Soma Biswas

机构 * Indian Institute of Science(印度科学研究院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2504.11002 2025-08-13 cs.SD cs.MM eess.AS 84%

Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation

Yan Rong, Shan Yang, Chenxing Li, Dong Yu, Li Liu

专题命中 音频语音多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01384 2025-08-13 cs.CV 83%

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing

Langyu Wang, Bingke Zhu, Yingying Chen, Yiyuan Zhang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, China(中国科学院自动化研究所基础模型研究中心)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accpted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03300 2025-08-13 cs.RO cs.LG 78%

Touch and Tell: Multimodal Decoding of Human Emotions and Social Gestures for Robots

Qiaoqiao Ren, Remko Proesmans, Yuanbo Hou, Francis wyffels, Tony Belpaeme

机构 * Faculty of Engineering and Architecture(工程与建筑学院) IDLab-AIRO, Ghent University – imec(IDLab-AIRO,根特大学–imec) Department of Engineering Science, University of Oxford(工程科学系,牛津大学)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15447 2025-08-13 cs.MM cs.CV cs.SD eess.AS 75%

Gotta Hear Them All: Towards Sound Source Aware Audio Generation

Wei Guo, Heng Wang, Jianbo Ma, Weidong Cai

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments 17 pages, 12 figures, source code available at https://github.com/wguo86/SSV2A

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2502.12454 2025-08-13 cs.CV cs.AI cs.HC cs.LG 82%

Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches

He Zhang, Xinyi Fu

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, accepted to MRAC'25: 3rd International Workshop on Multimodal and Responsible Affective Computing (ACM-MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02935 2025-08-13 cs.CL 79%

Dynamic Graph Neural ODE Network for Multi-modal Emotion Recognition in Conversation

Yuntao Shou, Tao Meng, Wei Ai, Keqin Li

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学教育部长江网络与网络安全重点实验室) College of Computer and Mathematics, Central South University of Forestry and Technology(中南林业科技大学计算机与数学学院) Department of Computer Science, State University of New York(纽约州立大学新帕尔茨分校计算机科学系)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08541 2025-08-13 physics.app-ph 78%

Multimodal learning enables instant ionizing radiation alerts on unmodified mobile phones for real-world emergency response

Yanfeng Xie, Xingzhi Cheng

专题命中 视频多模态 :multimodal(title,abstract)

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08989 2025-08-13 cs.CV 57%

KFFocus: Highlighting Keyframes for Enhanced Video Understanding

Ming Nie, Chunwei Wang, Hang Xu, Li Zhang

机构 * School of Data Science, Fudan University(复旦大学数据科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08590 2025-08-13 cs.CV cs.HC 57%

QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection

Yuxiao Wang, Wolin Liang, Yu Lei, Weiying Xue, Nan Zhuang, Qi Liu

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2508.08781 2025-08-13 cs.CV 74%

SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)

Trong-Thuan Nguyen, Viet-Tham Huynh, Quang-Thuc Nguyen, Hoang-Phuc Nguyen, Long Le Bao, Thai Hoang Minh, Minh Nguyen Anh, Thang Nguyen Tien, Phat Nguyen Thuan, Huy Nguyen Phong, Bao Huynh Thai, Vinh-Tiep Nguyen, Duc-Vu Nguyen, Phu-Hoa Pham, Minh-Huy Le-Hoang, Nguyen-Khang Le, Minh-Chinh Nguyen, Minh-Quan Ho, Ngoc-Long Tran, Hien-Long Le-Hoang, Man-Khoi Tran, Anh-Duong Tran, Kim Nguyen, Quan Nguyen Hung, Dat Phan Thanh, Hoang Tran Van, Tien Huynh Viet, Nhan Nguyen Viet Thien, Dinh-Khoi Vo, Van-Loc Nguyen, Trung-Nghia Le, Tam V. Nguyen, Minh-Triet Tran

机构 * University of Science, VNU-HCM(越南胡志明市科学大学) University of Information Technology, VNU-HCM(越南胡志明市信息技术大学) Vietnam National University(越南国家大学) University of Dayton(戴维森大学)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00589 2025-08-13 cs.CV cs.CL cs.IR cs.RO 62%

Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving

Stefan Englmeier, Max A. Büttner, Katharina Winter, Fabian B. Flohr

机构 * Munich University of Applied Sciences(慕尼黑应用科学大学) Intelligent Vehicles Lab (IVL)(智能车辆实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Project page: https://iv.ee.hm.edu/contextmotionclip/; This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2508.08821 2025-08-13 cs.CV 79%

3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs

Noor Ahmed, Cameron Braunstein, Steffen Eger, Eddy Ilg

专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11676 2025-08-13 cs.RO cs.AI cs.LG cs.MA 79%

Hypergraph-based Motion Generation with Multi-modal Interaction Relational Reasoning

Keshu Wu, Yang Zhou, Haotian Shi, Dominique Lord, Bin Ran, Xinyue Ye

机构 * organization= Center for Geospatial Sciences, Applications Department of Landscape of Architecture Urban Planning, Texas A\&M University , addressline= 788 Ross St , city= College Station , postcode= 77840 , state= TX , country= United States organization= Zachry Department of Civil Environmental Engineering, Texas A\&M University , addressline= 201 Dwight Look Engineering Building , city= College Station , postcode= 77843 , state= TX , country= United States organization= Department of Civil Environmental Engineering, University of Wisconsin-Madison , addressline= 1415 Engineering Dr , city= Madison , postcode= 53706 , state= WI , country= United States

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21349 2025-08-13 cs.CE 78%

Out of the Past: An AI-Enabled Pipeline for Traffic Simulation from Noisy, Multimodal Detector Data and Stakeholder Feedback

Rex Chen, Karen Wu, John McCartney, Norman Sadeh, Fei Fang

专题命中 多模态生成 :multimodal(title,abstract)

Comments 17 pages; 1 table; 6 figures; extended version of accepted version, published at the 2025 Winter Simulation Conference (WSC '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08987 2025-08-13 cs.CV cs.HC 74%

ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation

Ding Xia, Naoto Inoue, Qianru Qiu, Kotaro Kikuchi

机构 * The University of Tokyo(东京大学) CyberAgent AI Lab(CyberAgent AI实验室)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments Accepted to ICDAR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00043 2025-08-13 cs.CL cs.AI cs.CV 67%

CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation

Jixuan Leng, Chengsong Huang, Langlin Huang, Bill Yuchen Lin, William W. Cohen, Haohan Wang, Jiaxin Huang

机构 * CMU(卡内基梅隆大学) WUSTL(华盛顿大学) UIUC(伊利诺伊大学香槟分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12899 2025-08-13 cs.CV cs.AI cs.MM 67%

DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion

Huiguo He, Huan Yang, Zixi Tuo, Yuan Zhou, Qiuyue Wang, Yuhang Zhang, Zeyu Liu, Wenhao Huang, Hongyang Chao, Jian Yin

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08891 2025-08-13 cs.CV 57%

Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos

Chaoyi Wang, Yifan Yang, Jun Pei, Lijie Xia, Jianpo Liu, Xiaobing Yuan, Xinhan Di

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments This paper has been accepted by ICCV 2025 Workshop MMFM4

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09028 2025-08-13 cs.HC 50%

Envisioning Generative Artificial Intelligence in Cartography and Mapmaking

Yuhao Kang, Chenglong Wang

专题命中 多模态生成 :multimodal(abstract)

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 6 篇

2506.14805 2025-08-13 cs.CV cs.AI cs.CL cs.LG cs.MM 83%

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?

Yang Yao, Lingyu Li, Jiaxin Song, Chiyu Chen, Zhenqi He, Yixu Wang, Xin Wang, Tianle Gu, Jie Li, Yan Teng, Yingchun Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shanghai Jiao Tong University(上海交通大学) The Hong Kong University of Science and Technology(香港科学与技术大学) Fudan University(复旦大学) Tsinghua University(清华大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02141 2025-08-13 cs.CV cs.CL 81%

WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Jipeng Zhang, Xiangjian He, Song Wu, Xiaohan Xing, Sen Yang, Xiyue Wang, Linlin Shen

机构 * Shenzhen University(深圳大学) University of Nottingham Ningbo China(诺丁汉大学宁波分校) City University of Hong Kong(香港城市大学) Stanford University(斯坦福大学) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL

Comments ICCV 2025, 38 pages, 22 figures, 35 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01306 2025-08-13 cs.LG cs.CV 79%

ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation

Moran Yanuka, Morris Alper, Hadar Averbuch-Elor, Raja Giryes

机构 * Tel-Aviv University(特拉维夫大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ACL 2024 (Finding). For Project webpage, see https://moranyanuka.github.io/icc/

Journal ref Findings of the Association for Computational Linguistics: ACL 2024, pages 11048-11064, Bangkok, Thailand, August 2024

详情

展开后加载摘要…

URL PDF HTML 收藏