arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2508.04571 2025-08-07 cs.IR cs.CL cs.LG 83%

Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation

Claudio Pomo, Matteo Attimonelli, Danilo Danese, Fedelucio Narducci, Tommaso Di Noia

机构 * Sapienza University of Rome(罗马大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted as Full Research Papers at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03102 2025-08-06 cs.CV 83%

Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning

Tianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng Shi

机构 * Australian Institute for Machine Learning(澳大利亚机器学习研究所) The University of Adelaide(阿德莱德大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02525 2025-08-05 cs.AI 83%

Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model

Qifan Chen, Jin Cui, Cindy Duan, Yushuo Han, Yifei Shi

机构 * King’s College London(伦敦国王学院) Imperial College London(帝国理工学院) Columbia University(哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Submitted to the NeurIPS 2025 Workshop GenAI4Health. Conference website: https://aihealth.ischool.utexas.edu/GenAI4HealthNeurips2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01644 2025-08-05 cs.MM cs.AI cs.CV cs.SD eess.AS 83%

DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition

Peiyuan Jiang, Yao Liu, Qiao Liu, Zongshun Zhang, Jiaye Yang, Lu Liu, Daibing Yao

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Published in ACM Multimedia 2025. 10 pages, 4 figures

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23444 2025-08-01 cs.MM 83%

Hybrid CNN-Mamba Enhancement Network for Robust Multimodal Sentiment Analysis

Xiang Li, Xianfu Cheng, Xiaoming Zhang, Zhoujun Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08204 2025-07-30 cs.CV 83%

One-stage Modality Distillation for Incomplete Multimodal Learning

Shicai Wei, Yang Luo, Chunbo Luo

机构 * School of Information and Communication Engineering University of Electronic Science and Technology of China(信息与通信工程学院 电子科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19839 2025-07-29 cs.LG cs.CV 83%

GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning

Tiantian Peng, Yuyang Liu, Shuo Yang, Qiuhe Hong, YongHong Tian

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16158 2025-07-23 cs.CV 83%

AMMNet: An Asymmetric Multi-Modal Network for Remote Sensing Semantic Segmentation

Hui Ye, Haodong Chen, Zeke Zexi Hu, Xiaoming Chen, Yuk Ying Chung

机构 * School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学) School of Computer and Artificial Intelligence, Beijing Technology and Business University(计算机与人工智能学院,北京技术与商业大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15253 2025-07-22 cs.AI cs.LG cs.SI 83%

Disentangling Homophily and Heterophily in Multimodal Graph Clustering

Zhaochen Guo, Zhixiang Shen, Xuanting Xie, Liangjian Wen, Zhao Kang

机构 * University of Electronic Science and Technology of China(电子科技大学) Southwestern University of Finance and Economics(西南财经大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Appear in ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14935 2025-07-22 cs.CV 83%

Open-set Cross Modal Generalization via Multimodal Unified Representation

Hai Huang, Yan Xia, Shulei Wang, Hanting Wang, Minghui Fang, Shengpeng Ji, Sashuai Zhou, Tao Jin, Zhou Zhao

机构 * Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07886 2025-07-22 cs.CV 83%

EgoM2P: Egocentric Multimodal Multitask Pretraining

Gen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao, Marc Pollefeys, Siyu Tang

机构 * ETH Zürich(苏黎世联邦理工学院) Zhejiang University(浙江大学) Microsoft(微软公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17066 2025-07-17 cs.CV cs.LG 83%

DUNIA: Pixel-Sized Embeddings via Cross-Modal Alignment for Earth Observation Applications

Ibrahim Fayad, Max Zimmer, Martin Schwartz, Fabian Gieseke, Philippe Ciais, Gabriel Belouze, Sarah Brood, Aurelien De Truchis, Alexandre d'Aspremont

机构 * Laboratoire des Sciences du Climat et de l’Environnement, LSCE/IPSL, France Department for AI in Society, Science Technology, Zuse Institute Berlin, Germany Department of Information Systems, University of Münster, Germany Department of Computer Science, CNRS, INRIA \& École Normale Supérieure, Paris 75230, France

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 26 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10213 2025-07-15 cs.CV 83%

Boosting Multimodal Learning via Disentangled Gradient Learning

Shicai Wei, Chunbo Luo, Yang Luo

机构 * The Laboratory of Intelligent Collaborative Computing of UESTC(UESTC智能协同计算实验室) The School of Information and Communication Engineering of UESTC(UESTC信息与通信工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08855 2025-07-15 eess.IV cs.CV cs.LG 83%

Multi-omic Prognosis of Alzheimer's Disease with Asymmetric Cross-Modal Cross-Attention Network

Yang Ming, Jiang Shi Zhong, Zhou Su Juan

机构 * College of Medical Information Engineering, Guangdong Pharmaceutical University, Guangzhou, Guangdong 510006, China(医学信息工程学院,广东药科大学,广州,广东510006,中国)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07108 2025-07-11 cs.CV cs.AI cs.CL cs.LG cs.MM 83%

Multi-level Mixture of Experts for Multimodal Entity Linking

Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li, Jeff Z. Pan

机构 * School of Computer Information Technology Shanxi University Taiyuan China School of Computer Science Informatics Cardiff University Cardiff UK ILCC, School of Informatics University of Edinburgh Edinburgh UK Information Technology Shanxi University Informatics Cardiff University ILCC, School of Informatics University of Edinburgh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12151 2025-07-08 cs.LG cs.AI 83%

Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis

Miaosen Luo, Yuncheng Jiang, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04664 2025-07-08 cs.CV 83%

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

Tao Zhang, Shiqing Wei, Shihao Chen, Wenling Yu, Muying Luo, Shunping Ji

机构 * School of Remote Sensing and Information Engineering(遥感与信息工程学院) College of Oceanography and Space Informatics(海洋学与空间信息学院) School of Surveying and Geoinformation Engineering(测绘与地理信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04635 2025-07-08 cs.CV 83%

MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding

Zhicheng Zhang, Wuyou Xia, Chenxi Zhao, Zhou Yan, Xiaoqiang Liu, Yongjie Zhu, Wenyu Qin, Pengfei Wan, Di Zhang, Jufeng Yang

机构 * VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Pengcheng Laboratory

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICML 2025 (Spotlight, Top 2.6%)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03019 2025-07-08 cs.CV cs.LG 83%

Look-Back: Implicit Visual Re-focusing in MLLM Reasoning

Shuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye, Bin Lin, Li Yuan

机构 * Peking University(北京大学) Shenzhen Graduate School(深圳研究生院) Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02859 2025-07-04 cs.CV 83%

Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation

Jiaer Xia, Bingkui Tong, Yuhang Zang, Rui Shao, Kaiyang Zhou

机构 * Hong Kong Baptist University(香港 Baptist 大学) Sichuan University(四川大学) Shanghai AI Lab(上海人工智能实验室) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16236 2025-07-04 cs.CV 83%

LLaVA-KD: A Framework of Distilling Multimodal Large Language Models

Yuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He, Ao Tong, Zhenye Gan, Chengjie Wang, Zhucun Xue, Yong Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Zhejiang University(浙江大学) Youtu Lab, Tencent(腾讯优图实验室) Huazhong Agricultural University(华中农业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02080 2025-07-04 cs.MM cs.SD 83%

TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation

Yubeen Lee, Sangeun Lee, Chaewon Park, Junyeop Cha, Eunil Park

机构 * Sungkyunkwan University(全北大学) Electronics and Telecommunications Research Institute(电子电信研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments 9 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21096 2025-07-02 cs.CL 83%

DALR: Dual-level Alignment Learning for Multimodal Sentence Representation Learning

Kang He, Yuzhe Ding, Haining Wang, Fei Li, Chong Teng, Donghong Ji

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航天信息安全部门与可信计算重点实验室,教育部,网络安全科学与工程学院,武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted by ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23462 2025-07-01 cs.LG cs.AI 83%

Can We Predict the Unpredictable? Leveraging DisasterNet-LLM for Multimodal Disaster Classification

Manaswi Kulahara, Gautam Siddharth Kashyap, Nipun Joshi, Arpita Soni

机构 * TERI School Of Advanced Studies(TERI高级研究学院) Macquarie University(麦考瑞大学) Cornell University(康奈尔大学) Eudoxia Research University(欧多西亚研究大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Accepted in the 2025 IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2025), scheduled for 3 - 8 August 2025 in Brisbane, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22736 2025-07-01 cs.CV 83%

UniFuse: A Unified All-in-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and Misalignments

Dayong Su, Yafei Zhang, Huafeng Li, Jinxing Li, Yu Liu

机构 * Kunming University of Science and Technology(昆明理工大学) Harbin Institute of Technology at Shenzhen(哈尔滨工业大学深圳研究院) Hefei University of Technology(合肥工业大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19326 2025-07-01 cs.CV 83%

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment

Ziang Yan, Zhilin Li, Yinan He, Chenting Wang, Kunchang Li, Xinhao Li, Xiangyu Zeng, Zilei Wang, Yali Wang, Yu Qiao, Limin Wang, Yi Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Nanjing University(南京大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21536 2025-07-01 cs.CL 83%

Tracing Intricate Cues in Dialogue: Joint Graph Structure and Sentiment Dynamics for Multimodal Emotion Recognition

Jiang Li, Xiaoping Wang, Zhigang Zeng

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) Institute of Artificial Intelligence, Huazhong University of Science and Technology(华中科技大学人工智能研究院) Hubei Key Laboratory of Brain-Inspired Intelligent Systems, Huazhong University of Science and Technology(华中科技大学脑启发智能系统省重点实验室) Key Laboratory of Image Processing and Intelligent Control (Huazhong University of Science and Technology), Ministry of Education(图像处理与智能控制重点实验室(华中科技大学))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22446 2025-07-01 cs.LG cs.AI 83%

EAGLE: Efficient Alignment of Generalized Latent Embeddings for Multimodal Survival Prediction with Interpretable Attribution Analysis

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

机构 * Dept. of Machine Learning Moffitt Cancer Center(机器学习系莫菲特癌症中心) Dept. of Cancer Epidemiology Moffitt Cancer Center(癌症流行病学系莫菲特癌症中心) Dept. of Electrical Engineering University of South Florida(电气工程系佛罗里达州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13980 2025-06-25 cs.CV 83%

FusionSAM: Visual Multi-Modal Learning with Segment Anything

Daixun Li, Weiying Xie, Mingxiang Cao, Yunke Wang, Yusi Zhang, Leyuan Fang, Yunsong Li, Chang Xu

机构 * Xidian University(西安电子科技大学) University of Sydney(悉尼大学) Hunan University(湖南大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16804 2025-06-25 cs.LG cs.AI cs.CY cs.ET 83%

Multimodal Machine Learning in Mental Health: A Survey of Data, Algorithms, and Challenges

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * Queen Mary University of London United Kingdom Queen Mary University of London \& Jo z ef Stefan Institute United Kingdom \& Slovenia Queen Mary University of London Queen Mary University of London \& Jo z ef Stefan Institute

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏