arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-01 至 2025-09-01 共收录 22 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 2 篇

2312.15663 2025-09-01 cs.CV cs.AI 73%

IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models

Zhihao Chen, Bin Hu, Chuang Niu, Tao Chen, Yuxin Li, Hongming Shan, Ge Wang

机构 * ISTBI Fudan University(ISTBI 复旦大学) Huashan Hospital Fudan University(复旦大学华山医院) BME & CBIS Rensselaer Polytechnic Institute(生物医学工程与生物信息学研究中心罗切斯特理工学院)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 9 figures

Journal ref Visual Computing for Industry, Biomedicine, and Art, 7, 20, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21732 2025-09-01 cs.CV cs.AI 62%

CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models

João Valente, Atabak Dehban, Rodrigo Ventura

机构 * Institute for Systems and Robotics(系统与机器人研究所) University of Lisbon(里斯本大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2503.02823 2025-09-01 cs.SD cs.AI cs.MM eess.AS 82%

A Multimodal Symphony: Integrating Taste and Sound through Generative AI

Matteo Spanio, Massimiliano Zampini, Antonio Rodà, Franco Pierucci

机构 * Centro di Sonologia Computazionale (CSC)(计算声学中心) Department of Information Engineering University of Padova(信息工程系帕多瓦大学) Center for Mind/Brain Sciences (CIMeC)(心智/大脑科学中心) University of Trento(特伦托大学) SoundFood s.r.l.(SoundFood公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM、eess.AS

Comments 17 pages, 6 figures (2 + 2 figures with 2 subfigures each)

Journal ref Front. Comput. Sci. 7:1575741 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21761 2025-09-01 cs.CV cs.MM 73%

Learning from Silence and Noise for Visual Sound Source Localization

Xavier Juanola, Giovana Morais, Magdalena Fuentes, Gloria Haro

机构 * Intelligent Multimodal Vision Analysis Universitat Pompeu Fabra Barcelona, Spain(智能多模态视觉分析Universitat Pompeu Fabra Barcelona) MARL-IDM New York University, New York, USA(MARL-IDM纽约大学)

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.MM

Comments 10 pages, 2 figures, 4 tables + Supplementary Material

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11079 2025-09-01 eess.AS cs.AI cs.CL cs.SD 67%

Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts

Lingyun Gao, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments This paper is accepted to Interspeech 2025. This publication is part of the project Responsible AI for Voice Diagnostics (RAIVD) with file number NGF.1607.22.013 of the research programme NGF AiNed Fellowship Grants which is financed by the Dutch Research Council (NWO)

Journal ref https://www.isca-archive.org/interspeech_2025/gao25c_interspeech.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17625 2025-09-01 cs.LG cs.AI 57%

Alice's Adventures in a Differentiable Wonderland -- Volume I, A Tour of the Land

Simone Scardapane

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Companion website for additional chapters: https://www.sscardapane.it/alice-book

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 1 篇

2507.04061 2025-09-01 cs.CV cs.MM 73%

Consistent and Invariant Generalization Learning for Short-video Misinformation Detection

Hanghui Guo, Weijie Shi, Mengze Li, Juncheng Li, Hao Chen, Yue Cui, Jiajie Xu, Jia Zhu, Jiawei Shen, Zhangze Chen, Sirui Han

机构 * Zhejiang Normal University(浙江师范大学) Hong Kong University of Science and Technology(香港理工大学) Zhejiang University(浙江大学) Tencent(腾讯) Soochow University(苏州大学)

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 1 篇

2508.21539 2025-09-01 cs.CV 70%

HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones

Hao Ruan, Jinliang Lin, Yingxin Lai, Zhiming Luo, Shaozi Li

机构 * Department of Artificial Intelligence, Xiamen University(人工智能学院,厦门大学)

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 2 篇

2508.21460 2025-09-01 cs.IR cs.AI 79%

Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate Prediction

Xiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li, Zhejun Zhao

机构 * Peking University(北京大学) Wuhan University(武汉大学) Shanghai University of International Business(上海国际商务大学) Microsoft(微软)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI

Comments SIGIR 2025

Journal ref SIGIR 2025: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval Pages 581 - 591

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15695 2025-09-01 cs.LG 78%

SimuGen: Multi-modal Agentic Framework for Constructing Block Diagram-Based Simulation Models

Xinxing Ren, Qianbo Zang, Zekun Guo

机构 * Brunel University of London(伦敦布鲁内尔大学) SnT, Université du Luxembourg(卢森堡大学SnT分校) University of Hull(霍尔姆斯大学)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 7 篇

2508.21430 2025-09-01 cs.CL cs.AI cs.CV 85%

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models

Meidan Ding, Jipeng Zhang, Wenxuan Wang, Cheng-Yi Li, Wei-Chieh Fang, Hsin-Yu Wu, Haiqin Zhong, Wenting Chen, Linlin Shen

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室) The Hong Kong University of Science and Technology(香港科学与技术大学) Renmin University of China(中国人民大学) National Yang Ming Chiao Tung University Taipei Veterans General Hospital(台北荣民总医院) School of Biomedical Engineering, Shenzhen University(深圳大学生物医学工程学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 19 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12263 2025-09-01 cs.CV cs.AI 84%

Region-Level Context-Aware Multimodal Understanding

Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan, Debin Zhao

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院长) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21635 2025-09-01 cs.RO cs.CV cs.SY eess.SY 84%

The Rosario Dataset v2: Multimodal Dataset for Agricultural Robotics

Nicolas Soncini, Javier Cremona, Erica Vidal, Maximiliano García, Gastón Castro, Taihú Pire

机构 * CIFASIS (CONICET-UNR)(CIFASIS(CONICET-UNR)) Universidad de San Andrés (UDESA-CONICET)(Universidad de San Andrés(UDESA-CONICET))

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract,journal_ref);分类 cs.CV

Comments First published on The International Journal of Robotics Research: https://journals.sagepub.com/doi/10.1177/02783649251368909

Journal ref The Rosario dataset v2: Multi-modal dataset for agricultural robotics. The International Journal of Robotics Research. 2025;0(0)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08340 2025-09-01 cs.CV cs.AI 81%

Single Domain Generalization for Multimodal Cross-Cancer Prognosis via Dirac Rebalancer and Distribution Entanglement

Jia-Xuan Jiang, Jiashuai Liu, Hongtao Wu, Yifeng Wu, Zhong Wang, Qi Bi, Yefeng Zheng

机构 * Lanzhou University \& Westlake University Lanzhou China Xi'an Jiaotong University Xi'an China Westlake University \& Chinese University of Hong Kong University Hang Zhou, China Southern University of Science Lanzhou University Lanzhou China University of Amsterdam Amsterdam Netherland Westlake University Hangzhou China Lanzhou University \& Westlake University Xi'an Jiaotong University Westlake University \& Chinese University of Hong Kong University Lanzhou University University of Amsterdam Westlake University

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACMMM 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21793 2025-09-01 cs.LG cs.AI 79%

MoE-Health: A Mixture of Experts Framework for Robust Multimodal Healthcare Prediction

Xiaoyang Wang, Christopher C. Yang

机构 * Drexel University(德雷塞尔大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to The 16th ACM Conference on Bioinformatics, Computational Biology, and Health Informatics (ACM-BCB 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21169 2025-09-01 cs.CV 79%

SYNBUILD-3D: A large, multi-modal, and semantically rich synthetic dataset of 3D building models at Level of Detail 4

Kevin Mayer, Alex Vesel, Xinyi Zhao, Martin Fischer

机构 * Department of Civil and Environmental Engineering, Stanford University(土木与环境工程系,斯坦福大学) Department of Computer Science, Stanford University(计算机科学系,斯坦福大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21579 2025-09-01 cs.CR 50%

Agentic Discovery and Validation of Android App Vulnerabilities

Ziyue Wang, Liyi Zhou

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 2 篇

2508.21364 2025-09-01 cs.RO cs.SY eess.SY 78%

Multi-Modal Model Predictive Path Integral Control for Collision Avoidance

Alberto Bertipaglia, Dariu M. Gavrila, Barys Shyrokau

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 多模态Agent :multi-modal(title,abstract)

Comments Accepted as an oral presentation at the 29th IAVSD. August 18-22, 2025. Shanghai, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21456 2025-09-01 cs.HC cs.CL cs.CV 62%

Morae: Proactively Pausing UI Agents for User Choices

Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, Amy Pavel

机构 * Carnegie Mellon University(卡内基梅隆大学) Adobe Research(Adobe研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

Comments ACM UIST 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 多模态训练与对齐 2 篇

2508.21581 2025-09-01 cs.CV 61%

Integrating Pathology and CT Imaging for Personalized Recurrence Risk Prediction in Renal Cancer

Daniël Boeke, Cedrik Blommestijn, Rebecca N. Wray, Kalina Chupetlovska, Shangqi Gao, Zeyu Gao, Regina G. H. Beets-Tan, Mireia Crispin-Ortuzar, James O. Jones, Wilson Silva, Ines P. Machado

机构 * Department of Radiology, Antoni van Leeuwenhoek-Netherlands Cancer Institute, Amsterdam, The Netherlands(放射科,Antoni van Leeuwenhoek荷兰癌症研究所,阿姆斯特丹,荷兰) University of Amsterdam(阿姆斯特丹大学) AI Technology for Life, Department of Information and Computing Sciences, Department of Biology, Utrecht University(AI技术生命,信息与计算科学系,生物学系,乌得勒支大学) GROW Oncology, Maastricht University(GROW肿瘤学,马斯特里赫特大学) Department of Oncology, University of Cambridge(肿瘤科,剑桥大学) Cancer Research UK Cambridge Centre, University of Cambridge(英国癌症研究会剑桥中心,剑桥大学) Early Cancer Institute, University of Cambridge(早期癌症研究所,剑桥大学) Cambridge University Hospitals NHS Foundation Trust(剑桥大学医院 NHS 基础信托)

专题命中 多模态训练与对齐 :multimodal(abstract,comments);分类 cs.CV

Comments 12 pages, 2 figures, 1 table. Accepted at the Multimodal Learning and Fusion Across Scales for Clinical Decision Support (ML-CDS) Workshop, MICCAI 2025. This is the submitted version with authors, affiliations, and acknowledgements included; it has not undergone peer review or revisions. The final version will appear in the Springer Lecture Notes in Computer Science (LNCS) proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21824 2025-09-01 cs.CV 57%

DriveQA: Passing the Driving Knowledge Test

Maolin Wei, Wanzhou Liu, Eshed Ohn-Bar

机构 * Boston University(波士顿大学) Washington University in St. Louis(华盛顿大学圣路易斯分校)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025. Project page: https://driveqaiccv.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

9. 其他多模态 1 篇

2508.21801 2025-09-01 cs.IR 82%

DMGIN: How Multimodal LLMs Enhance Large Recommendation Models for Lifelong User Post-click Behaviors

Zhuoxing Wei, Qingchen Xie, Qi Liu

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract)

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏