arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-28 至 2025-08-28 共收录 44 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 3 篇

2411.19930 2025-08-28 cs.CL cs.CV cs.LG 84%

On Domain-Adaptive Post-Training for Multimodal Large Language Models

Daixuan Cheng, Shaohan Huang, Ziyu Zhu, Xintong Zhang, Wayne Xin Zhao, Zhongzhi Luan, Bo Dai, Zhenliang Zhang

机构 * BIGAI BUAA(北京航空航天大学) THU(清华大学) BIT(北京理工大学) RUC(中国人民大学)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract,comments);分类 cs.CV、cs.CL

Comments EMNLP 2025 Findings, Project Page: https://huggingface.co/AdaptLLM/Adapt-MLLM-to-Domains

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18634 2025-08-28 cs.CV 70%

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward

Chunlin Zhong, Qiuxia Hou, Zhangjun Zhou, Shuang Hao, Haonan Lu, Yanhao Zhang, He Tang, Xiang Bai

专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments 9 pages, 6figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19376 2025-08-28 cs.LG cs.AI cs.CV hep-ex 62%

Fine-Tuning Vision-Language Models for Neutrino Event Analysis in High-Energy Physics Experiments

Dikshant Sagar, Kaiwen Yu, Alejandro Yankelevich, Jianming Bian, Pierre Baldi

机构 * Department of Computer Science University of California, Irvine(计算机科学系加州大学伊文斯顿分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 5 篇

2408.16564 2025-08-28 cs.MM cs.SD eess.AS 84%

Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition

Qianhui Liu, Jiadong Wang, Yang Wang, Xin Yang, Gang Pan, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.MM、eess.AS

Comments aceepted by IEEE TC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19483 2025-08-28 eess.AS 79%

Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids

Nasir Saleem, Mandar Gogate, Kia Dashtipour, Adeel Hussain, Usman Anwar, Adewale Adetomi, Tughrul Arslan, Amir Hussain

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Preprint of the paper presented at Euronoise 2025 Malaga, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16632 2025-08-28 cs.CL cs.SD eess.AS 62%

Step-Audio 2 Technical Report

Boyong Wu, Chao Yan, Chen Hu, Cheng Yi, Chengli Feng, Fei Tian, Feiyu Shen, Gang Yu, Haoyang Zhang, Jingbei Li, Mingrui Chen, Peng Liu, Wang You, Xiangyu Tony Zhang, Xingyuan Li, Xuerui Yang, Yayue Deng, Yechang Huang, Yuxin Li, Yuxin Zhang, Zhao You, Brian Li, Changyi Wan, Hanpeng Hu, Jiangjie Zhen, Siyu Chen, Song Yuan, Xuelin Zhang, Yimin Jiang, Yu Zhou, Yuxiang Yang, Bingxin Li, Buyun Ma, Changhe Song, Dongqing Pang, Guoqiang Hu, Haiyang Sun, Kang An, Na Wang, Shuli Gao, Wei Ji, Wen Li, Wen Sun, Xuan Wen, Yong Ren, Yuankai Ma, Yufan Lu, Bin Wang, Bo Li, Changxin Miao, Che Liu, Chen Xu, Dapeng Shi, Dingyuan Hu, Donghang Wu, Enle Liu, Guanzhe Huang, Gulin Yan, Han Zhang, Hao Nie, Haonan Jia, Hongyu Zhou, Jianjian Sun, Jiaoren Wu, Jie Wu, Jie Yang, Jin Yang, Junzhe Lin, Kaixiang Li, Lei Yang, Liying Shi, Li Zhou, Longlong Gu, Ming Li, Mingliang Li, Mingxiao Li, Nan Wu, Qi Han, Qinyuan Tan, Shaoliang Pang, Shengjie Fan, Siqi Liu, Tiancheng Cao, Wanying Lu, Wenqing He, Wuxun Xie, Xu Zhao, Xueqi Li, Yanbo Yu, Yang Yang, Yi Liu, Yifan Lu, Yilei Wang, Yuanhao Ding, Yuanwei Liang, Yuanwei Lu, Yuchu Luo, Yuhe Yin, Yumeng Zhan, Yuxiang Zhang, Zidong Yang, Zixin Zhang, Binxing Jiao, Daxin Jiang, Heung-Yeung Shum, Jiansheng Chen, Jing Li, Xiangyu Zhang, Yibo Zhu

机构 * StepFun Audio Team(StepFun音频团队)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

Comments v3: Added introduction and evaluation results of Step-Audio 2 mini

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08400 2025-08-28 cs.CL cs.LG cs.SD eess.AS 62%

mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks

Luel Hagos Beyene, Vivek Verma, Min Ma, Jesujoba O. Alabi, Fabian David Schmidt, Joyce Nakatumba-Nabende, David Ifeoluwa Adelani

机构 * AIMS RIC NM-AIST Mila - Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Google DeepMind(谷歌DeepMind) Saarland University(萨尔兰大学) University of Würzburg(维尔茨堡大学) Makerere University(Makerere大学) McGill University(麦吉尔大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted to ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19514 2025-08-28 cs.SD 50%

MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models

Zhihao Ouyang, Ju-Chiang Wang, Daiyu Zhang, Bin Chen, Shangjie Li, Quan Lin

机构 * ByteDance(字节跳动)

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 3 篇

2508.19567 2025-08-28 cs.LG 78%

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning

Sheryl Mathew, N Harshit

专题命中 视频多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19639 2025-08-28 cs.MM 57%

FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter

Junxi Wang, Yaxiong Wang, Lechao Cheng, Zhun Zhong

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

Comments EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19916 2025-08-28 cond-mat.mtrl-sci 50%

Microscale optoelectronic reservoir networks of halide perovskite for in-sensor computing

Jeroen J. de Boer, Agustin O. Alvarez, Moritz C. Schmidt, Bruno Ehrler

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 5 篇

2508.20057 2025-08-28 cs.MM 83%

ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation

Haoshuo Zhang, Yufei Bo, Meixia Tao

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments arXiv admin note: text overlap with arXiv:2508.17920

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19319 2025-08-28 eess.IV cs.AI cs.CV 81%

MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction

Pardis Moradbeiki, Nasser Ghadiri, Sayed Jalal Zahabi, Uffe Kock Wiil, Kristoffer Kittelmann Brockhattingen, Ali Ebrahimi

机构 * Department of Electrical and Computer Engineering, Isfahan University of Technology(电气与计算机工程系,伊斯法罕技术大学) SDU Health Informatics and Technology, The Maersk Mc-Kinney Moller Institute, University of Southern Denmark(南部丹麦大学健康信息学与技术,马士基麦金尼莫勒研究所) Geriatric Research Unit, Department of Clinical Research, University of Southern Denmark(老年医学研究单元,临床研究系,南部丹麦大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07460 2025-08-28 cs.LG cs.AI cs.DB 79%

HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19596 2025-08-28 cs.AI cs.IR 57%

Reference-Aligned Retrieval-Augmented Question Answering over Heterogeneous Proprietary Documents

Nayoung Choi, Grace Byun, Andrew Chung, Ellie S. Paek, Shinsun Lee, Jinho D. Choi

机构 * Department of Computer Science Emory University Atlanta Georgia USA(计算机科学系 埃默里大学 阿拉巴马 州 美国) Emory University(埃默里大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments Accepted to CIKM 2025 Applied Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19942 2025-08-28 cs.HC 50%

Socially Interactive Agents for Preserving and Transferring Tacit Knowledge in Organizations

Martin Benderoth, Patrick Gebhard, Christian Keller, C. Benjamin Nakhosteen, Stefan Schaffer, Tanja Schneeberger

专题命中 跨模态检索 :multimodal(abstract)

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 7 篇

2502.09242 2025-08-28 cs.AI 79%

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

Lukas Buess, Matthias Keicher, Nassir Navab, Andreas Maier, Soroosh Tayebi Arasteh

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

Journal ref Biomed. Eng. Lett. 15 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20020 2025-08-28 cs.CV 57%

GS: Generative Segmentation via Label Diffusion

Yuhao Chen, Shubin Chen, Liang Lin, Guangrun Wang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 12 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19508 2025-08-28 cs.RO cs.CV 57%

DATR: Diffusion-based 3D Apple Tree Reconstruction Framework with Sparse-View

Tian Qiu, Alan Zoubi, Yiyuan Lin, Ruiming Du, Lailiang Cheng, Yu Jiang

机构 * School of Electrical and Computer Engineering, Cornell University(电气与计算机工程系,康奈尔大学) Sibley School of Mechanical and Aerospace Engineering, Cornell University(机械与航空航天工程系,康奈尔大学) School of Biological and Environmental Engineering, Cornell University(生物与环境工程系,康奈尔大学) School of Integrative Plant Science, Cornell University(整合植物科学系,康奈尔大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15842 2025-08-28 cs.CV cs.GR 57%

DiffArtist: Towards Structure and Appearance Controllable Image Stylization

Ruixiang Jiang, Changwen Chen

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACM MM 2025, Homepage: https://DiffusionArtist.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14874 2025-08-28 cs.CV 57%

TraceNet: Segment one thing efficiently

Mingyuan Wu, Zichuan Liu, Haozhen Zheng, Hongpeng Guo, Bo Chen, Xin Lu, Klara Nahrstedt

机构 * Coordinated Science Laboratory, University of Illinois at Urbana-Champaign, Champaign, USA(伊利诺伊大学厄巴纳-香槟分校协调科学实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Best Student Paper in IEEE MIPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12166 2025-08-28 cs.RO cs.LG cs.SY eess.SY 50%

Belief-Conditioned One-Step Diffusion: Real-Time Trajectory Planning with Just-Enough Sensing

Gokul Puthumanaillam, Aditya Penumarti, Manav Vora, Paulo Padrao, Jose Fuentes, Leonardo Bobadilla, Jane Shin, Melkior Ornik

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Florida(佛罗里达大学) Providence College(普罗维登斯学院) Florida International University(佛罗里达国际大学)

专题命中 多模态生成 :multi-modal(abstract)

Comments Accepted to CoRL 2025 (Conference on Robot Learning)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14207 2025-08-28 cs.RO 50%

A Comprehensive Review on Traffic Datasets and Simulators for Autonomous Vehicles

Supriya Sarker, Brent Maples, Iftekharul Islam, Muyang Fan, Christos Papadopoulos, Weizi Li

专题命中 多模态生成 :multimodal(abstract)

Comments This manuscript has been withdrawn due to the need for substantial updates and revisions

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 10 篇

2508.20068 2025-08-28 cs.CL cs.CV cs.LG 84%

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis

Chengzu Li, Wenshan Wu, Huanyu Zhang, Qingtao Li, Zeyu Gao, Yan Xia, José Hernández-Orallo, Ivan Vulić, Furu Wei

机构 * Microsoft Research(微软研究院) Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) Department of Oncology, University of Cambridge(癌症部门,剑桥大学) Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能中心,剑桥大学) VRAIN, Universitat Politècnica de València(VRAIN,巴塞罗那理工大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments 9 pages, 4 figures (22 pages, 7 figures, 7 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01490 2025-08-28 q-bio.GN cs.AI cs.CV cs.LG q-bio.TO stat.AP 84%

A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics

Rushin H. Gindra, Giovanni Palla, Mathias Nguyen, Sophia J. Wagner, Manuel Tran, Fabian J Theis, Dieter Saur, Lorin Crawford, Tingying Peng

机构 * Helmholtz Munich, Germany(海德堡-慕尼黑赫尔姆霍尔茨研究中心) Technical University Munich, Germany(慕尼黑技术大学) Microsoft Research, USA(微软研究院)

专题命中 多模态评测 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

Comments The code is accessible at: https://github.com/peng-lab/hescape

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19659 2025-08-28 cs.LG 82%

SCAR: A Characterization Scheme for Multi-Modal Dataset

Ri Su, Zhao Chen, Caleb Chen Cao, Nan Tang, Lei Chen

机构 * HKUST (GZ)(香港科技大学(珠海)) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract)

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16439 2025-08-28 cs.CY cs.AI cs.CL cs.GR cs.MM 82%

PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark

Adil Bahaj, Oumaima Fadi, Mohamed Chetouani, Mounir Ghogho

机构 * Mohammed 6 Polytechnic University(摩洛哥6号理工学院) International University of Rabat(拉巴特国际大学) Institut des Systèmes Intelligents et de Robotique(智能系统与机器人研究所)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04836 2025-08-28 eess.IV cs.AI cs.CV 81%

PGAD: Prototype-Guided Adaptive Distillation for Multi-Modal Learning in AD Diagnosis

Yanfei Li, Teng Yin, Wenyi Shang, Jingyu Liu, Xi Wang, Kaiyang Zhao

机构 * Machine Intelligence Lab, College of Computer Science and Technology(机器智能实验室,计算机科学与技术学院) Sichuan University(四川大学) Department of Computer Science and Engineering(计算机科学与工程系) The Chinese University of Hong Kong(香港中文大学) Department of Neurosurgery(神经外科部门)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19862 2025-08-28 cs.CV cs.LG 79%

Multimodal Conditional MeshGAN for Personalized Aneurysm Growth Prediction

Long Chen, Ashiv Patel, Mengyun Qiao, Mohammad Yousuf Salmasi, Salah A. Hammouche, Vasilis Stavrinides, Jasleen Nagi, Soodeh Kalaie, Xiao Yun Xu, Wenjia Bai, Declan P. O'Regan

机构 * MRC Laboratory of Medical Sciences Imperial College London(医学科学实验室 Imperial College London) Imperial College Healthcare NHS Trust(帝国理工医疗 NHS Trust) Department of Mechanical Engineering University College London(机械工程系 University College London) London Postgraduate School of Surgery NHS England(伦敦外科研究生学校 NHS England) Faculty of Medicine Imperial College London(医学系 Imperial College London) Department of Chemical Engineering Imperial College London(化学工程系 Imperial College London) Department of Brain Sciences&Computing Imperial College London(脑科学与计算系 Imperial College London)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18088 2025-08-28 cs.RO cs.AI cs.CL cs.CV cs.MA 67%

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, Weiliang Deng, Yubin Guo, Tian Nian, Xuanbing Xie, Qiangyu Chen, Kailun Su, Tianling Xu, Guodong Liu, Mengkang Hu, Huan-ang Gao, Kaixuan Wang, Zhixuan Liang, Yusen Qin, Xiaokang Yang, Ping Luo, Yao Mu

机构 * MoE key Lab of Artificial Intelligence, AI Institute, SJTU(人工智能联合实验室、人工智能研究院、上海交通大学) HKU MMLab(香港大学多模态实验室) Shanghai AI Lab(上海人工智能实验室) D-Robotics(D-机器人) SZU(深圳大学) THU(清华大学) TeleAI FDU(福建大学) USTC(中国科学技术大学) SUSTech(南方科技大学) SYSU(深圳大学) CSU(中国科学技术大学) NEU(南京大学) HKU-SH ICRC(香港大学深圳研究院) NJU(南京大学) Lumina EAI

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Project Page: https://robotwin-platform.github.io/, Code: https://github.com/robotwin-Platform/robotwin, Doc: https://robotwin-platform.github.io/doc/

详情

展开后加载摘要…

URL PDF HTML 收藏