arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2602.04920 2026-02-06 cs.LG cs.SD 82%

CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning

CyIN:循环信息潜在空间用于连接完整与不完整多模态学习

Ronghao Lin, Qiaolin He, Sijie Mai, Ying Zeng, Aolin Xiong, Li Huang, Yap-Peng Tan, Haifeng Hu

机构 * School of Electronics and Information Technology, Sun Yat-Sen University(中山大学电子与信息学院) School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电气与电子工程学院) School of Computer Science, South China Normal University(华南师范大学计算机科学学院) Desay SV Automotive Co., Ltd(德赛股份有限公司) Pazhou Laboratory(琶洲实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 CyIN通过构建循环信息潜在空间,解决多模态学习中完整与不完整数据之间的性能差距,实现统一优化。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04231 2026-02-05 cs.RO 82%

GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning

GeoLanG: 基于几何的语言引导抓取与统一RGB-D多模态学习

Rui Tang, Guankun Wang, Long Bai, Huxin Gao, Jiewen Lai, Chi Kit Ng, Jiazheng Wang, Fan Zhang, Hongliang Ren

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 GeoLanG通过统一RGB-D多模态学习,结合深度引导几何模块和自适应通道集成,实现鲁棒的语言引导抓取,提升复杂环境中的抓取精度和泛化能力。

Comments IEEE ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04016 2026-02-05 eess.SP cs.LG 82%

A Multi-Modal Foundational Model for Wireless Communication and Sensing

一种用于无线通信和传感的多模态基础模型

Vahid Yazdnian, Yasaman Ghasempour

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 本文提出了一种多模态基础模型,通过物理指导的自监督预训练策略,实现无线通信和传感任务的稳健泛化与高效适应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21436 2026-02-05 cs.LG cs.AI cs.CL cs.CV 82%

From Consistency to Complementarity: Aligned and Disentangled Multi-modal Learning for Time Series Understanding and Reasoning

从一致性到互补性:面向时间序列理解和推理的对齐与解缠多模态学习

Hang Ni, Weijia Zhang, Fei Wang, Zezhi Shao, Hao Liu

机构 * The Hong Kong University of Science(香港科学与技术大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 MADI通过细粒度对齐和解缠交互提升多模态时间序列理解和推理能力,实现更精确的数值-视觉模态整合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00914 2026-02-03 cs.CL cs.AI cs.CY cs.SD eess.AS 82%

A Baseline Multimodal Approach to Emotion Recognition in Conversations

一种用于对话中情感识别的基线多模态方法

Víctor Yeste, Rodrigo Rivas-Arévalo

机构 * School of Science, Engineering and Design, Universidad Europea de Valencia(科学、工程与设计学院,欧洲大学 Valencia)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本文提出了一种基于Transformer文本分类器和自监督语音模型的多模态基线方法,用于对话中情感识别,并通过实验展示了多模态融合的优势。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15388 2026-02-03 cs.CV cs.AI cs.CL 82%

LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models

LLaVA-PruMerge: 适应性令牌减少用于高效的大多模态模型

Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, Yan Yan

机构 * UCF(佛罗里达大学) UW-Madison(威斯康星大学麦迪逊分校) USC(南加州大学) UIC(伊利诺伊大学香槟分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 LLaVA-PruMerge通过自适应视觉令牌减少策略,显著降低视觉令牌数量而不影响性能,适用于高效的大多模态模型。

Comments Accepted to ICCV 2025. First Version is released in 2024/03

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22498 2026-02-02 cs.IR 82%

FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation Learning

FITMM: 一种基于信息论的多模态推荐方法

Wei Yang, Rui Zhong, Yiqun Chen, Shixuan Li, Heng Ping, Chi Lu, Peng Jiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 FITMM通过频域信息论框架提升多模态推荐效果,采用频谱分解与信息瓶颈目标实现频带分离与融合,有效减少冗余并提升泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17809 2026-01-27 cs.IT math.IT 82%

A Multi-Modal Fusion Platform for Joint Environment Sensing and Channel Sounding in Highly Dynamic Scenarios

一种多模态融合平台,用于高动态场景中的联合环境感知与信道探测

Xuejian Zhang, Ruisi He, Mi Yang, Zhengyu Zhang, Ziyi Qi

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)

AI总结 本文提出一种多模态融合平台,用于高动态场景中联合环境感知与信道探测,支持多频段多天线测量和高精度环境感知。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15304 2026-01-27 cs.IR 82%

MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation

MLLMRec: 基于图细化的多模态推荐偏好推理范式

Yuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan, Wanyu Chen, Fei Cai, Honghui Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract)

AI总结 MLLMRec通过图细化和多模态大语言模型提升多模态推荐的用户偏好推理与物品表示学习准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00670 2026-01-05 cs.HC 82%

Wave2Word: A Multimodal Transformer Framework for Joint EEG-Text Alignment and Multi-Task Representation Learning in Neurocritical Care

Wave2Word: 一种多模态Transformer框架,用于神经重症监护中的联合EEG-文本对齐和多任务表示学习

Argha Kamal Samanta, Deepak Mewada, Monalisa Sarma, Debasis Samanta

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 Wave2Word提出一种多模态Transformer框架,通过整合信号域建模与结构化临床语言监督,实现EEG-文本对齐和多任务表示学习,提升神经重症监护中的EEG分析效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08270 2025-12-30 cs.LG cs.AI cs.CL cs.MM 82%

Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI

Doctor Sun: 一种双语多模态大语言模型用于生物医学AI

Dong Xue, Ziyao Shao, Zhaoyang Duan, Fangzhou Liu, Bing Li, Zhongheng Zhang

机构 * Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学) Research Institute of Intelligent Control and Systems Harbin Institute of Technology(智能控制与系统研究室,哈尔滨工业大学) Department of Emergency Medicine, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(浙江大学医学院急诊医学科) Provincial Key Laboratory of Precise Diagnosis Treatment of Abdominal Infection, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(腹部感染精准诊断治疗省级重点实验室,浙江大学医学院) School of Medicine Shaoxing University(绍兴大学医学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 Doctor Sun是一种双语多模态大语言模型,通过整合预训练视觉编码器和医学LLM,提升生物医学多模态任务的性能,并提供SunMed-VL数据集支持研究进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18686 2025-12-22 cs.LG 82%

Hierarchical Multimodal LLMs with Semantic Space Alignment for Enhanced Time Series Classification

具有语义空间对齐的层次多模态大语言模型用于增强的时间序列分类

Xiaoyu Tao, Tingyue Pan, Mingyue Cheng, Yucong Luo, Qi Liu, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 HiTime通过层次多模态大语言模型和语义空间对齐,提升时间序列分类的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22410 2025-12-09 stat.AP 82%

Multimodal Fusion and Interpretability in Human Activity Recognition: A Reproducible Framework for Sensor-Based Modeling

多模态融合与可解释性在人体活动识别中的应用:一种可复现的基于传感器建模框架

Yiyao Yang, Yasemin Gulbahar

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract)

AI总结 本文提出了一种可复现的多模态融合框架,通过统一预处理和融合策略提升人体活动识别的准确性和可解释性。

Comments 33 pages, 12 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16990 2025-11-24 cs.HC 82%

Senti-iFusion: An Integrity-centered Hierarchical Fusion Framework for Multimodal Sentiment Analysis under Uncertain Modality Missingness

Senti-iFusion: 一种以完整性为中心的多模态情感分析多模态融合框架,用于在不确定模态缺失情况下

Liling Li, Guoyang Xu, Xiongri Shen, Zhifei Xu, Yanbo Zhang, Zhiguo Zhang, Zhenxi Song

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 Senti-iFusion提出了一种以完整性为中心的多模态融合框架,通过分层结构处理模态缺失问题,提升多模态情感分析的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00374 2025-11-18 cs.CV cs.AI cs.MM 82%

MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention

Tianyi Wang, Jianan Fan, Dingxin Zhang, Dongnan Liu, Yong Xia, Heng Huang, Weidong Cai

机构 * The University of Sydney(悉尼大学) School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) School of Computer Science and Engineering, Northwestern Polytechnical University(西北工业大学计算机科学与工程学院) University of Maryland(马里兰大学) Ningbo Institute of Northwestern Polytechnical University(西北工业大学宁波学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by IEEE Transactions on Medical Imaging (TMI). Code available at https://github.com/TianyiFranklinWang/MIRROR. Project page: https://tianyifranklinwang.github.io/MIRROR

Journal ref IEEE Trans. Med. Imaging (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10568 2025-11-17 cs.LG cs.AI cs.CL cs.CV 82%

MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion

Ruixiang Jiang, Lingbo Liu, Changwen Chen

机构 * Department of Computing, The Hong Kong Polytechnic University(计算系,香港理工大学) Research Institute of Multiple Agents and Embodied Intelligence, Pengcheng Laboratory(多智能体与具身智能研究院,鹏城实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to IEEE TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07274 2025-11-11 cs.LG 82%

Multi-modal Dynamic Proxy Learning for Personalized Multiple Clustering

Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Ziyue Peng, Zewei Liu, Hewei Wang, Jiayi Zhang, Edith C. H. Ngai

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03371 2025-11-06 cond-mat.mtrl-sci physics.comp-ph 82%

Enhancing composition-based materials property prediction by cross-modal knowledge transfer

Ivan Rubtsov, Ivan Dudakov, Yuri Kuratov, Vadim Korolev

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

Comments 7 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07841 2025-11-04 cs.NI cs.LG 82%

Task-Oriented Multimodal Token Transmission in Resource-Constrained Multiuser Networks

Junhe Zhang, Wanli Ni, Pengwei Wang, Dongyu Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00987 2025-11-04 cs.LG 82%

Balanced Multimodal Learning via Mutual Information

Rongrong Xie, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati (SISSA)(国际先进研究学院(SISSA))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11194 2025-10-27 cs.CE 82%

Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive Alignment

Xiao Fei, Michail Chatzianastasis, Sarah Almeida Carneiro, Hadi Abdine, Lawrence P. Petalidis, Michalis Vazirgiannis

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 24 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20736 2025-10-24 cs.LG 82%

Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process

Tsai Hor Chan, Feng Wu, Yihang Chen, Guosheng Yin, Lequan Yu

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by NeruIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20540 2025-10-24 cs.LG 82%

SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment

Abdulmomen Ghalkha, Zhuojun Tian, Chaouki Ben Issaid, Mehdi Bennis

机构 * Center for Wireless Communications, University of Oulu(无线通信中心,奥卢大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments 5 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20256 2025-10-24 cs.CV cs.CL cs.LG cs.MM 82%

Calibrating Multimodal Consensus for Emotion Recognition

Guowei Zhong, Junjie Li, Huaiyu Zhu, Ruohong Huan, Yun Pan

机构 * College of Information Science and Electronic Engineering, Zhejiang University(信息科学与电子工程学院,浙江大学) College of Computer Science and Technology, Zhejiang University of Technology(计算机科学与技术学院,浙江工业大学) Zhejiang University Jinhua Research Institute(浙江大学金华研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16824 2025-10-21 cs.LG q-bio.MN 82%

ProtoMol: Enhancing Molecular Property Prediction via Prototype-Guided Multimodal Learning

Yingxu Wang, Kunyu Zhang, Jiaxin Huang, Nan Yin, Siwei Liu, Eran Segal

机构 * MBZUAI(穆扎芬人工智能研究所) University of Zhengzhou(郑州大学) HKUST(香港科技大学) University of Aberdeen(爱丁堡大学) MBZUAI, Weizmann Institute of Science(穆扎芬人工智能研究所、威斯曼科学研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16350 2025-10-21 cs.LG 82%

MGTS-Net: Exploring Graph-Enhanced Multimodal Fusion for Augmented Time Series Forecasting

Shule Hao, Junpeng Bao, Wenli Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12395 2025-10-15 cs.CR 82%

IP-Augmented Multi-Modal Malicious URL Detection Via Token-Contrastive Representation Enhancement and Multi-Granularity Fusion

Ye Tian, Yanqiu Yu, Liangliang Song, Zhiquan Liu, Yanbin Wang, Jianguo Sun

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15595 2025-10-14 cs.RO 82%

Grasping Deformable Objects via Reinforcement Learning with Cross-Modal Attention to Visuo-Tactile Inputs

Yonghyun Lee, Sungeun Hong, Min-gu Kim, Gyeonghwan Kim, Changjoo Nam

机构 * Dept. of Electronic Engineering at Sogang University(ソガン大学电子工程系) Dept. of Immersive Media and Engineering at Sungkyunkwan University(顺天大学沉浸媒体与工程系) College of Medicine, Yonsei University(延世大学医学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08022 2025-10-10 cs.LG cs.AI cs.CL cs.CV 82%

Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining

Chenxi Liu, Tianyi Xiong, Yanshuo Chen, Ruibo Chen, Yihan Wu, Junfeng Guo, Tianyi Zhou, Heng Huang

机构 * University of Maryland, College Park(马里兰大学学院 park)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05283 2025-10-08 cs.AI cs.CL cs.CV 82%

Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment

Radha Gulhane, Sathish Reddy Indurthi

机构 * Radha Gulhane(独立研究者) Sathish Reddy Indurthi(独立研究者)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏