arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-22 至 2025-09-22 共收录 63 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 11 篇

2501.18592 2025-09-22 cs.CV cs.AI cs.LG cs.RO 84%

Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models

Hao Dong, Moru Liu, Kaiyang Zhou, Eleni Chatzi, Juho Kannala, Cyrill Stachniss, Olga Fink

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

Comments Project page: https://github.com/donghao51/Awesome-Multimodal-Adaptation

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05244 2025-09-22 cs.CV cs.AI 84%

RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding

Tianchen Fang, Guiru Liu

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Upon further review, we identified that our dataset requires optimization to ensure research reliability and accuracy. Additionally, considering the target journal's latest submission policies, we believe comprehensive manuscript revisions are necessary

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15243 2025-09-22 cs.CV 84%

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

Muhammad Imran, Yugyung Lee

机构 * Computer Science, School of Science and Engineering, University of Missouri - Kansas City(计算机科学系,科学与工程学院,密苏里大学-堪萨斯城分校)

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV;multimodal(journal_ref)

Comments 8 pages, 6 figures, 3 tables

Journal ref Non-Archival track - The First Workshop on Multimodal Knowledge and Language Modeling IJCAI 2025 Workshop, August 16, 2025 IJCAI 2025 Workshop, August 16, 2025 Room 516B, Palais des congrès, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16163 2025-09-22 cs.CV cs.AI cs.CL 69%

Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks

Het Patel, Muzammil Allie, Qian Zhang, Jia Chen, Evangelos E. Papalexakis

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 图文多模态 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI

Comments To be presented as a poster at the Workshop on Safe and Trustworthy Multimodal AI Systems (SafeMM-AI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05439 2025-09-22 cs.CV cs.AI cs.CL 67%

LLMs Can Compensate for Deficiencies in Visual Representations

Sho Takishita, Jay Gala, Abdelrahman Mohamed, Kentaro Inui, Yova Kementchedjhieva

机构 * Fujitsu Limited(富士通有限公司) MBZUAI Tohoku University(东北大学) RIKEN(日本研究机构)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15490 2025-09-22 cs.CV cs.AI 62%

SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters

Abdarahmane Traore, Éric Hervet, Andy Couturier

机构 * Embia, Computer Science Department, Faculty of Science, Université de Moncton(Embia计算机科学系,科学学院,蒙特龙大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 3 figures, IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14312 2025-09-22 cs.CV 57%

CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation

Marc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosier, Nicolas Thome

机构 * Conservatoire National des Arts et Métiers(法国国家艺术与工艺学院) Sorbonne Université(索邦大学) ETS Montreal(蒙特利尔ETS) Institut universitaire de France(法国国家科学研究中心)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Journal ref 39th Conference on Neural Information Processing Systems, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07620 2025-09-22 cs.CV 57%

ViLU: Learning Vision-Language Uncertainties for Failure Prediction

Marc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon, Clément Rambour, Raphaël Fournier-Sniehotta, Ismail Ben Ayed, Jose Dolz, Nicolas Thome

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Journal ref International Conference on Computer Vision, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00743 2025-09-22 cs.CV 57%

Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models

Dilxat Muhtar, Enzhuo Zhang, Zhenshi Li, Feng Gu, Yanglangxing He, Pengfeng Xiao, Xueliang Zhang

机构 * Nanjing University(南京大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 39 pages, 13 figures. Accept for NeruIPS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14033 2025-09-22 cs.CV 57%

SAIL-VL2 Technical Report

Weijie Yin, Yongjie Ye, Fangxun Shu, Yue Liao, Zijian Kang, Hongyuan Dong, Haiyang Yu, Dingkang Yang, Jiacong Wang, Han Wang, Wenzhuo Liu, Xiao Liang, Shuicheng Yan, Chao Feng

机构 * Douyin SAIL Team(字节跳动 SAIL 团队) LV-NUS Lab(NUS 实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14967 2025-09-22 cs.RO cs.HC 50%

Affordance-Based Disambiguation of Surgical Instructions for Collaborative Robot-Assisted Surgery

Ana Davila, Jacinto Colan, Yasuhisa Hasegawa

机构 * Nagoya University, Japan(名古屋大学)

专题命中 图文多模态 :multimodal(abstract)

Comments To be presented at the 1st Workshop on Intelligent Cobodied Assistance and Robotic Empowerment (iCARE). 2025 Conference on Robot Learning (CoRL)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 6 篇

2509.16193 2025-09-22 eess.AS 89%

Are Multimodal Foundation Models All That Is Needed for Emofake Detection?

Mohd Mujtaba Akhtar, Girish, Orchid Chetia Phukan, Swarup Ranjan Behera, Pailla Balakrishna Reddy, Ananda Chandra Nayak, Sanjib Kumar Nayak, Arun Balaji Buduru

专题命中 音频语音多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract);分类 eess.AS

Comments Accepted to APSIPA-ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16025 2025-09-22 cs.CL cs.AI 88%

Session-Level Spoken Language Assessment with a Multimodal Foundation Model via Multi-Target Learning

Hong-Yun Lin, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen

机构 * Department of Computer Science(计算机科学系) Information Engineering, National Taiwan Normal University(信息工程,台湾正常大学)

专题命中 音频语音多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CL、cs.AI

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15476 2025-09-22 cs.CL cs.MM 86%

Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding

Zhu Li, Xiyuan Gao, Yuqing Zhang, Shekhar Nayak, Matt Coler

机构 * University of Groningen, The Netherlands(Groningen大学,荷兰)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15661 2025-09-22 cs.SD cs.AI cs.CL eess.AS 85%

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

Qiaolin Wang, Xilin Jiang, Linyang He, Junkai Wu, Nima Mesgarani

机构 * Columbia University(哥伦比亚大学) University of Washington(华盛顿大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16023 2025-09-22 eess.AS 79%

Interpreting the Role of Visemes in Audio-Visual Speech Recognition

Aristeidis Papadopoulos, Naomi Harte

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted into Automatic Speech Recognition and Understanding- ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15775 2025-09-22 cs.SD eess.AS 70%

EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model

Yiqing Yang, Man-Wai Mak

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子与电气工程系,香港理工大学)

专题命中 音频语音多模态 :multimodal(abstract);MLLM(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 6 篇

2509.16087 2025-09-22 cs.CV cs.AI 84%

See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model

Pengteng Li, Pinhao Song, Wuyang Li, Weiyu Guo, Huizai Yao, Yijie Xu, Dugang Liu, Hui Xiong

机构 * HKUST(GZ)(香港科技大学(广州)) KU Leuven(比利时鲁文大学) EPFL(苏黎世联邦理工学院) SZU(深圳大学)

专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15400 2025-09-22 cs.LG cs.AI cs.RO 79%

Exploring multimodal implicit behavior learning for vehicle navigation in simulated cities

Eric Aislan Antonelo, Gustavo Claudio Karl Couto, Christian Möller

机构 * Systems Engineering Department, Federal University of Santa Catarina, Florianopolis, Brazil Faculty of Science Engineering, Information Technology Åbo Akademi University, Finland

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments ENIAC conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08578 2025-09-22 cs.LG q-bio.PE q-bio.QM 78%

Multi-modal Adaptive Estimation for Temporal Respiratory Disease Outbreak

Hong Liu, Kerui Cen, Yanxing Chen, Zige Liu, Dong Chen, Zifeng Yang, Chitin Hon

机构 * Respiratory Disease AI Laboratory in Epidemic Intelligence and Applications of Medical Big Data Instruments, Macau University of Science and Technology(呼吸疾病人工智能实验室(流行病智能与医学大数据应用)) Faculty of Innovation Engineering, Macau University of Science and Technology(创新工程学院) Institute of Systems Engineering, Macau University of Science and Technology(系统工程研究所) School of Business, Macau University of Science and Technology(商学院) State Key Laboratory of Respiratory Disease, National Clinical Research Center for Respiratory Disease, Guangzhou Institute of Respiratory Health, The First Affiliated Hospital of Guangzhou Medical University(呼吸疾病国家重点实验室、呼吸疾病临床研究中心、广州呼吸健康研究院、广州医学院第一附属医院) Guangzhou National Laboratory(广州国家实验室)

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15233 2025-09-22 cs.MM cs.CL cs.CV 78%

Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents

Xueqiao Zhang, Chao Zhang, Jingtao Xu, Yifan Zhu, Xin Shi, Yi Yang, Yawei Luo

机构 * Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.MM

Comments Accepted at EMNLP2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15448 2025-09-22 cs.LG cs.AI cs.NE stat.ML 57%

Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems

Saeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito Koishida

机构 * Microsoft Redmond, WA 98052(微软红mond分校)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments In The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15481 2025-09-22 cs.LG cs.SI 50%

Solar Forecasting with Causality: A Graph-Transformer Approach to Spatiotemporal Dependencies

Yanan Niu, Demetri Psaltis, Christophe Moser, Luisa Lambertini

机构 * EPFL(苏黎世联邦理工学院)

专题命中 视频多模态 :multimodal(abstract)

Comments Accepted to CIKM 2025

Journal ref Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25), November 10--14, 2025, Seoul, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2509.15882 2025-09-22 cs.CV cs.AI 81%

Self-Supervised Cross-Modal Learning for Image-to-Point Cloud Registration

Xingmei Wang, Xiaoyu Hu, Chengkai Huang, Ziyan Zeng, Guohao Nie, Quan Z. Sheng, Lina Yao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15470 2025-09-22 cs.CV cs.AI 81%

Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture

Thomas Z. Li, Aravind R. Krishnan, Lianrui Zuo, John M. Still, Kim L. Sandler, Fabien Maldonado, Thomas A. Lasko, Bennett A. Landman

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14031 2025-09-22 cs.NE cs.LG 50%

Modeling the Human Visual System: Comparative Insights from Response-Optimized and Task-Optimized Vision Models, Language Models, and different Readout Mechanisms

Shreya Saha, Ishaan Chadha, Meenakshi Khosla

机构 * Electrical and Computer Engineering University of California, San Diego(电气与计算机工程大学加州大学圣地亚哥分校) Halıcıoğlu Data Science Institute University of California, San Diego(Halıcıoğlu数据科学研究所大学加州大学圣地亚哥分校) Department of Cognitive Science, Department of Computer Science and Engineering University of California, San Diego(认知科学系计算机科学与工程系大学加州大学圣地亚哥分校)

专题命中 跨模态检索 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2509.16127 2025-09-22 cs.CV 83%

BaseReward: A Strong Baseline for Multimodal Reward Model

Yi-Fan Zhang, Haihua Yang, Huanyu Zhang, Yang Shi, Zezhou Chen, Haochen Tian, Chaoyou Fu, Haotian Wang, Kai Wu, Bo Cui, Xu Wang, Jianfei Pan, Haotian Wang, Zhang Zhang, Liang Wang

机构 * ByteDance(字节跳动) CASIA(中国科学院自动化研究所) NJU(南京大学) PKU(北京大学) THU(清华大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13759 2025-09-22 cs.LG cs.AI 83%

Discrete Diffusion in Large Language and Multimodal Models: A Survey

Runpeng Yu, Qi Li, Xinchao Wang

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16197 2025-09-22 cs.CV cs.CL cs.LG 81%

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

Yanghao Li, Rui Qian, Bowen Pan, Haotian Zhang, Haoshuo Huang, Bowen Zhang, Jialing Tong, Haoxuan You, Xianzhi Du, Zhe Gan, Hyunjik Kim, Chao Jia, Zhenbang Wang, Yinfei Yang, Mingfei Gao, Zi-Yi Dou, Wenze Hu, Chang Gao, Dongxu Li, Philipp Dufter, Zirui Wang, Guoli Yin, Zhengdong Zhang, Chen Chen, Yang Zhao, Ruoming Pang, Zhifeng Chen

机构 * Apple(苹果公司)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15553 2025-09-22 cs.CV cs.AI stat.AP 81%

Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification

Tian Lan, Yiming Zheng, Jianxin Yin

机构 * School of Statistics, Renmin University of China(中国人民大学统计学院) Center for Applied Statistics and School of Statistics, Renmin University of China(中国人民大学应用统计中心和统计学院)

专题命中 多模态生成 :cross-modal(title);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏