arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-27 至 2025-08-27 共收录 44 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 2 篇

2405.11793 2025-08-27 cs.CV 83%

MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise

Ruiqi Wu, Chenran Zhang, Jianle Zhang, Yi Zhou, Tao Zhou, Huazhu Fu

机构 * School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, China(教育部新一代人工智能技术及其交叉应用重点实验室) Nanjing University of Science and Technology, Nanjing, China(南京理工大学) Agency for Science, Technology and Research (A*STAR), Singapore(新加坡科技研究局)

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Early Accepted by The International Conference on Medical Image Computing and Computer Assisted Intervention(MICCAI)2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13152 2025-08-27 cs.CV cs.AI cs.RO 81%

SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models

Xiangyu Dong, Haoran Zhao, Jiang Gao, Haozhou Li, Xiaoguang Ma, Yaoming Zhou, Fuhai Chen, Juan Liu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2508.18734 2025-08-27 cs.CV cs.AI cs.MM eess.AS eess.SP 89%

Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion

DongHoon Lim, YoungChae Kim, Dong-Hyun Kim, Da-Hee Yang, Joon-Hyuk Chang

机构 * Dept. Artificial Intelligence(人工智能系) Hanyang University(翰阳大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to IEEE ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18653 2025-08-27 cs.LG cs.AI cs.SD eess.AS 81%

The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability

Xiaoliang Chen, Xin Yu, Le Chang, Teng Jing, Jiashuai He, Ze Wang, Yangjun Luo, Xingyu Chen, Jiayue Liang, Yuchen Wang, Jiaying Xie

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22863 2025-08-27 cs.HC cs.CL 57%

Large Language Models for Depression Recognition in Spoken Language Integrating Psychological Knowledge

Yupei Li, Shuaijie Shao, Manuel Milling, Björn W. Schuller

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Journal ref Frontiers in Computer Science, Volume 7, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2504.17163 2025-08-27 cs.CV 83%

PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion Recognition

Kai Cui, Jia Li, Yu Liu, Xuesong Zhang, Zhenzhen Hu, Meng Wang

机构 * School of Instrument Science and Opto-electronics Engineering, Hefei University of Technology(仪器科学与光电工程学院,合肥工业大学) School of Computer Science and Information Engineering, Hefei University of Technology(计算机科学与信息工程学院,合肥工业大学)

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments To appear in IEEE TCSS. The source code is publicly available at https://github.com/MSA-LMC/PhysioSync

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18284 2025-08-27 cs.LG cs.AI cs.SY eess.SY 79%

Multi-Modal Drift Forecasting of Leeway Objects via Navier-Stokes-Guided CNN and Sequence-to-Sequence Attention-Based Models

Rahmat K. Adesunkanmi, Alexander W. Brandt, Masoud Deylami, Gustavo A. Giraldo Echeverri, Hamidreza Karbasian, Adel Alaeddini

机构 * Departments of Mechanical and Electrical Engineering, Southern Methodist University(机械与电气工程系,南方 Methodist 大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18922 2025-08-27 cs.LG cs.AI 57%

HierCVAE: Hierarchical Attention-Driven Conditional Variational Autoencoders for Multi-Scale Temporal Modeling

Yao Wu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18904 2025-08-27 cs.CV 57%

Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025

Thien-Phuc Tran, Minh-Quang Nguyen, Minh-Triet Tran, Tam V. Nguyen, Trong-Le Do, Duy-Nam Ly, Viet-Tham Huynh, Khanh-Duy Le, Mai-Khiem Tran, Trung-Nghia Le

机构 * University of Science(科学大学) University of Dayton(代顿大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18725 2025-08-27 cs.NI cs.IT math.IT 50%

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions

Ruichen Zhang, Guangyuan Liu, Yinqiu Liu, Changyuan Zhao, Jiacheng Wang, Yunting Xu, Dusit Niyato, Jiawen Kang, Yonghui Li, Shiwen Mao, Sumei Sun, Xuemin Shen, Dong In Kim

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2507.02826 2025-08-27 cs.CV 83%

Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach

Panpan Ji, Junni Song, Yifan Lu, Hang Xiao, Hanyu Liu, Chao Li

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24073 2025-08-27 cs.AI cs.CL cs.CV 82%

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation

Chan-Wei Hu, Yueqi Wang, Shuo Xing, Chia-Ju Chen, Suofei Feng, Ryan Rossi, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯农工大学) University of California, Berkeley(加州大学伯克利分校) University of Texas at Austin(德克萨斯大学奥斯汀分校) Stanford University(斯坦福大学) Adobe Research(Adobe研究院)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18772 2025-08-27 cs.CV cs.CL 73%

Beyond the Textual: Generating Coherent Visual Options for MCQs

Wanqiang Wang, Longzhu He, Wei Zheng

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19097 2025-08-27 cs.AI 57%

Reasoning LLMs in the Medical Domain: A Literature Survey

Armin Berger, Sarthak Khanna, David Berghaus, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫研究所媒体工程部门) University of Bonn - Department of Computer Science(波恩大学计算机科学系) West-AI - Federal Ministry of Education and Research(西德人工智能 - 教育与研究部)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14940 2025-08-27 cs.LG 50%

Cohort-Aware Agents for Individualized Lung Cancer Risk Prediction Using a Retrieval-Augmented Model Selection Framework

Chongyu Qu, Allen J. Luna, Thomas Z. Li, Junchao Zhu, Junlin Guo, Juming Xiong, Kim L. Sandler, Bennett A. Landman, Yuankai Huo

机构 * Vanderbilt University(范德堡大学) Vanderbilt University Medical Center(范德堡大学医学中心)

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 6 篇

2508.11433 2025-08-27 cs.CV 83%

MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation

Qian Liang, Yujia Wu, Kuncheng Li, Jiwei Wei, Shiyuan He, Jinyu Guo, Ning Xie

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15194 2025-08-27 cs.CV cs.AI cs.LG 81%

DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models

Sungnyun Kim, Junsoo Lee, Kibeom Hong, Daesik Kim, Namhyuk Ahn

机构 * KAIST AI(韩国科学技术院人工智能研究所) NAVER WEBTOON AI Sookmyung Women’s University(成均馆女子大学) Inha University(釜山大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Expert Systems with Applications 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18421 2025-08-27 cs.CV 70%

Why Relational Graphs Will Save the Next Generation of Vision Foundation Models?

Fatemeh Ziaeetabar

机构 * Department of Computer Science, School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran, Tehran, Iran(塔里斯坦大学科学学院计算机科学系)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15761 2025-08-27 cs.CV 57%

Waver: Wave Your Way to Lifelike Video Generation

Yifu Zhang, Hao Yang, Yuqi Zhang, Yifei Hu, Fengda Zhu, Chuang Lin, Xiaofeng Mei, Yi Jiang, Bingyue Peng, Zehuan Yuan

机构 * Bytedance Waver Team(字节跳动Waver团队)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06905 2025-08-27 cs.CV 57%

MultiRef: Controllable Image Generation with Multiple Visual References

Ruoxi Chen, Dongping Chen, Siyuan Wu, Sinan Wang, Shiyun Lang, Petr Sushko, Gaoyang Jiang, Yao Wan, Ranjay Krishna

机构 * Zhejiang Wanli University(浙江万里大学) University of Washington(华盛顿大学) Huazhong University of Science and Technology(华中科技大学) Allen Institute for AI(人工智能研究院)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

Comments Accepted to ACM MM 2025 Datasets

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11639 2025-08-27 cs.LG 50%

Deep Generative Methods and Tire Architecture Design

Fouad Oubari, Raphael Meunier, Rodrigue Décatoire, Mathilde Mougeot

机构 * ENS Paris-Saclay, Centre Borelli(巴黎-萨克雷大学ENS分校,Borelli中心) Université Paris-Saclay, CNRS, ENS Paris-Saclay, Centre Borelli(巴黎-萨克雷大学,法国国家科学研究中心,巴黎-萨克雷大学ENS分校,Borelli中心) ENSIIE(Évry法国的ENSIEE)

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 8 篇

2508.18740 2025-08-27 cs.CL cs.AI 81%

M3HG: Multimodal, Multi-scale, and Multi-type Node Heterogeneous Graph for Emotion Cause Triplet Extraction in Conversations

Qiao Liang, Ying Shen, Tiantian Chen, Lin Zhang

机构 * Tongji University(同济大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 16 pages, 8 figures. Accepted to Findings of ACL 2025

Journal ref Findings of ACL 2025 (2025) 11416-11431

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18608 2025-08-27 cs.AI 79%

eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases

Janet Wang, Xin Hu, Yunbei Zhang, Diabate Almamy, Vagamon Bamba, Konan Amos Sébastien Koffi, Yao Koffi Aubin, Zhengming Ding, Jihun Hamm, Rie R. Yotsu

机构 * Tulane University(Tulane大学) Bakke Graduate University(Bakke研究生大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18506 2025-08-27 cs.CV 79%

DoGFlow: Self-Supervised LiDAR Scene Flow via Cross-Modal Doppler Guidance

Ajinkya Khoche, Qingwen Zhang, Yixi Cai, Sina Sharif Mansouri, Patric Jensfelt

机构 * KTH Royal Institute of Technology(皇家理工学院) Autonomous Transport Solutions Lab, Scania Group(自主运输解决方案实验室,斯堪尼亚集团)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00152 2025-08-27 cs.CL 79%

Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data

Ekaterina Borisova, Fabio Barth, Nils Feldhus, Raia Abu Ahmad, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Sebastian Möller

机构 * Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI)(德国人工智能研究中心有限公司) Technische Universität Berlin(柏林技术大学) BIFOLD Deutsche Telekom(德国电信) Common Crawl Foundation(Common Crawl基金会) Humboldt-Universität zu Berlin(柏林洪堡大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments TRL@ACL 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18372 2025-08-27 cs.CV 79%

OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding

Hieu Nguyen, Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science(科学大学) University of Dayton(戴维森大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07155 2025-08-27 cs.CV 79%

Meta-Learned Modality-Weighted Knowledge Distillation for Robust Multi-Modal Learning with Missing Data

Hu Wang, Salma Hassan, Yuyuan Liu, Congbo Ma, Yuanhong Chen, Qing Li, Jiahui Geng, Bingjie Wang, Yu Tian, Yutong Xie, Jodie Avery, Louise Hull, Ian Reid, Mohammad Yaqub, Gustavo Carneiro

机构 * University of Adelaide, Australia(澳大利亚阿德莱德大学) New York University Abu Dhabi, UAE(阿联酋纽约大学阿布扎克分校) University of Surrey, UK(英国萨里大学) Boai hospital of Zhongshan, China(中国中山博爱医院) University of Central Florida, USA(美国佛罗里达大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18430 2025-08-27 cs.CV cs.AI 62%

CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering

Aranya Saha, Tanvir Ahmed Khan, Ismam Nur Swapnil, Mohammad Ariful Haque

机构 * Aranya Saha ∗ , Tanvir Ahmed Khan ∗ , Ismam Nur Swapnil ∗ , and Mohammad Ariful Haque ∗(无明确机构)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 8 figures, Prepared for submission to IEEE Transactions on Human-Machine Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14137 2025-08-27 cs.CV cs.CL 62%

VAGUE: Visual Contexts Clarify Ambiguous Expressions

Heejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung, Youngjae Yu

机构 * Brown University(布朗大学) UC Berkeley(加州大学伯克利分校) Yonsei University(延世大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

Comments ICCV 2025, 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 4 篇

2412.18292 2025-08-27 cs.RO 78%

Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration

Zhixuan Shen, Haonan Luo, Kexun Chen, Fengmao Lv, Tianrui Li

机构 * Zhixuan Shen, Haonan Luo, Kexun Chen, Fengmao Lv, Tianrui Li(作者)

专题命中 多模态Agent :multimodal(title,abstract)

Comments 16 pages, 10 figures, Extended Version of accepted AAAI 2025 Paper

详情

展开后加载摘要…

URL PDF HTML 收藏