arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2507.04758 2025-09-18 cs.MM 79%

Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation Learning

Jiayun Hu, Yueyi He, Tianyi Liang, Changbo Wang, Chenhui Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12901 2025-09-17 cs.CV 79%

MSGFusion: Multimodal Scene Graph-Guided Infrared and Visible Image Fusion

Guihui Li, Bowei Dong, Kaizhi Dong, Jiayi Li, Haiyong Zheng

机构 * College of Computer Science and Technology, Ocean University of China(中国海洋大学计算机科学与技术学院) College of Electronic Engineering, Ocean University of China(中国海洋大学电子工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10986 2025-09-16 cs.CV cs.RO 79%

Long-Tailed 3D Detection via Multi-Modal Fusion

Yechi Ma, Neehar Peri, Achal Dave, Wei Hua, Deva Ramanan, Shu Kong

机构 * Department of Computer Science(计算机科学系) Robotics Institute(机器人研究所) Toyota Research Institute(丰田研究机构) Faculty of Science and Technology(科学与技术学院) Institute of Collaborative Innovation(协同创新研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments The first two authors contributed equally. Project page: https://mayechi.github.io/lt3d-lf-io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11376 2025-09-16 cs.LG cs.AI cs.CE 79%

Intelligent Reservoir Decision Support: An Integrated Framework Combining Large Language Models, Advanced Prompt Engineering, and Multimodal Data Fusion for Real-Time Petroleum Operations

Seyed Kourosh Mahjour, Seyed Saman Mahjour

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17910 2025-09-15 cs.LG cs.AI 79%

A Novel Approach to Balance Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes and its Implementation in BEACON

Vansh Nagpal, Siva Likitha Valluru, Kausik Lakkaraju, Nitin Gupta, Zach Abdulrahman, Andrew Davison, Biplav Srivastava

机构 * BEACON

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09064 2025-09-12 cs.CV 79%

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

Qiuhui Chen, Xuancheng Yao, Huping Ye, Yi Hong

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Journal of Biomedical and Health Informatics (JBHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04256 2025-09-11 cs.CV 79%

Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation

Zifu Wan, Pingping Zhang, Yuhao Wang, Silong Yong, Simon Stepputtis, Katia Sycara, Yaqi Xie

机构 * Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by WACV 2025. Project page: https://zifuwan.github.io/Sigma/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05615 2025-09-09 cs.LG cs.AI 79%

Causal Debiasing Medical Multimodal Representation Learning with Missing Modalities

Xiaoguang Zhu, Lianlong Sun, Yang Liu, Pengyi Jiang, Uma Srivatsa, Nipavan Chiamvimonvat, Vladimir Filkov

机构 * DataLab: Data Science and Informatics, University of California, Davis(加州大学戴维斯分校数据实验室) Department of Electrical and Computer Engineering, University of Rochester(罗切斯特大学电气与计算机工程系) Academy for Engineering & Technology, Fudan University(复旦大学工程与技术学院) Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Department of Electrical and Computer Engineering, New York University(纽约大学电气与计算机工程系) UC Davis Health(加州大学戴维斯分校医疗中心) Department of Basic Medical Sciences, University of Arizona(亚利桑那大学基础医学系) Department of Computer Science, University of California, Davis(加州大学戴维斯分校计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Submitted to IEEE TKDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04480 2025-09-08 cs.CL cs.LG 79%

Discrete Prompt Tuning via Recursive Utilization of Black-box Multimodal Large Language Model for Personalized Visual Emotion Recognition

Ryo Takahashi, Naoki Saito, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama

机构 * Graduate School of Information Science(信息科学研究生学校) Technology, Hokkaido University, Sapporo 060-0814, Japan(技术,北海道大学,札幌060-0814,日本) Office of Institutional Research, Hokkaido University, Sapporo 060-0808, Japan(机构研究办公室,北海道大学,札幌060-0808,日本) Faculty of Information Science(信息科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03999 2025-09-05 cs.CV 79%

SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation

Han Huang, Han Sun, Ningzhong Liu, Huiyu Zhou, Jiaquan Shen

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Leicester(莱斯特大学) Luoyang Normal University(洛阳师范学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages, accepted by PRCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03408 2025-09-04 cs.CV cs.LG 79%

Scalable and Loosely-Coupled Multimodal Deep Learning for Breast Cancer Subtyping

Mohammed Amer, Mohamed A. Suliman, Tu Bui, Nuria Garcia, Serban Georgescu

机构 * Fujitsu Research of Europe Ltd(富士通欧洲研究有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06594 2025-09-03 q-fin.CP cs.AI cs.LG 79%

Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism

Chang Zong, Hang Zhou

机构 * School of Information and Electronic Engineering, Zhejiang University of Science and Technology(浙江理工大学信息与电子工程学院) Department of Finance Accounting and Economics, Business School of Nottingham University(诺丁汉大学商学院金融会计与经济学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18057 2025-09-03 cs.SD cs.AI 79%

Dynamic Fusion Multimodal Network for SpeechWellness Detection

Wenqiang Sun, Han Yin, Jisheng Bai, Jianfeng Chen

机构 * Northwestern Polytechnical University(西北工业大学) School of Electrical Engineering, KAIST(韩国成均馆大学电气工程学院) LianFeng Acoustic Technologies Co., Ltd.(联丰声学科技有限公司) Xi’an University of Posts & Telecommunications(西安邮电大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 6 pages, 5figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00677 2025-09-03 cs.CV 79%

CSFMamba: Cross State Fusion Mamba Operator for Multimodal Remote Sensing Image Classification

Qingyu Wang, Xue Jiang, Guozheng Xu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 5 pages, 2 figures, accpeted by 2025 IEEE International Geoscience and Remote Sensing Symposium(IGARSS 2025),not published yet

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20415 2025-08-29 cs.CV 79%

Graph-Based Uncertainty Modeling and Multimodal Fusion for Salient Object Detection

Yuqi Xiong, Wuzhen Shi, Yang Wen, Ruhan Liu

机构 * Guangdong Key Laboratory of Intelligent Information Processing(广东智能信息处理重点实验室) College of Electronics and Information Engineering(电子信息工程学院) Shenzhen University(深圳大学) Furong Laboratory(芙蓉实验室) Central South University(中南大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ICONIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18050 2025-08-26 cs.CV 79%

ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation

Jianwen Tan, Huiyao Zhang, Rui Xiong, Han Zhou, Hongfei Wang, Ye Li

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17478 2025-08-26 cs.CV 79%

GraphMMP: A Graph Neural Network Model with Mutual Information and Global Fusion for Multimodal Medical Prognosis

Xuhao Shan, Ruiquan Ge, Jikui Liu, Linglong Wu, Chi Zhang, Siqi Liu, Wenjian Qin, Wenwen Min, Ahmed Elazab, Changmiao Wang

机构 * Hangzhou Dianzi University(杭州电子科技大学) Shenzhen Polytechnic University(深圳职业技术大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Yunnan University(云南大学) Shenzhen University(深圳大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17213 2025-08-26 cs.CV 79%

Multi-modal Knowledge Decomposition based Online Distillation for Biomarker Prediction in Breast Cancer Histopathology

Qibin Zhang, Xinyu Hao, Qiao Chen, Rui Xu, Fengyu Cong, Cheng Lu, Hongming Xu

机构 * School of Biomedical Engineering, Faulty of Medicine, Dalian University of Technology, Dalian, China(生物医学工程学院) Faculty of Information Technology, University of Jyvaskyla, Jyvaskyla, Finland(信息科技学院) School of Software Technology, Dalian University of Technology, Dalian, China(软件技术学院) Department of Radiology, Guangdong Provincial People’s Hospital, Southern Medical University, Guangzhou, China(放射科)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16408 2025-08-25 cs.CV 79%

SAMFusion: Sensor-Adaptive Multimodal Fusion for 3D Object Detection in Adverse Weather

Edoardo Palladin, Roland Dietze, Praveen Narayanan, Mario Bijelic, Felix Heide

机构 * Torc Robotics(Torc机器人公司) University of Stuttgart(斯图加特大学) Princeton University(普林斯顿大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16000 2025-08-25 eess.IV cs.CV cs.LG 79%

Cross-Attention Multimodal Fusion for Breast Cancer Diagnosis: Integrating Mammography and Clinical Data with Explainability

Muhaisin Tiyumba Nantogmah, Abdul-Barik Alhassan, Salamudeen Alhassan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15852 2025-08-25 cs.LG cs.CL 79%

PGF-Net: A Progressive Gated-Fusion Framework for Efficient Multimodal Sentiment Analysis

Bin Wen, Tien-Ping Tan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15505 2025-08-22 cs.CV 79%

Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion

Mengyu Wang, Zhenyu Liu, Kun Li, Yu Wang, Yuwei Wang, Yanyan Wei, Fei Wang

机构 * Key Laboratory of Opto-Electronic Information Science and Technology of Jiangxi Province, Nanchang Hangkong University(江西省光电信息科学与技术重点实验室,南昌航空大学) ReLER, CCAI, Zhejiang University(ReLER、CCAI、浙江大学) College of Engineering, Anhui Agricultural University(安徽农业大学工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13792 2025-08-22 cs.LG cs.AI cs.RO 79%

Continual Learning for Multimodal Data Fusion of a Soft Gripper

Nilay Kushawaha, Egidio Falotico

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted in Wiley Advanced Robotics Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12992 2025-08-20 cs.MM 79%

MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex Scenarios

Xiangxian Li, Yawen Zheng, Baiqiao Zhang, Yijia Ma, Xianhui Cao, Juan Liu, Yulong Bian, Jin Huang, Chenglei Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13196 2025-08-20 cs.LG cs.AI cs.IR 79%

Contextual Attention-Based Multimodal Fusion of LLM and CNN for Sentiment Analysis

Meriem Zerkouk, Miloud Mihoubi, Belkacem Chikhaoui

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments The 38th Canadian Conference on Artificial Intelligence ( 2025 )

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11350 2025-08-18 cs.CV 79%

HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model

Zhenhao Zhang, Hanqing Wang, Xiangyu Zeng, Ziyu Cheng, Jiaxin Liu, Haoyu Yan, Zhirui Liu, Kaiyang Ji, Tianxiang Gui, Ke Hu, Kangyi Chen, Yahao Fan, Mokai Pan

专题命中 多模态训练与对齐 :multimodal(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09717 2025-08-14 cs.CV cs.LG 79%

Multimodal Sheaf-based Network for Glioblastoma Molecular Subtype Prediction

Shekhnaz Idrissova, Islem Rekik

机构 * BASIRA Lab, Imperial-X(BASIRA实验室、Imperial-X) Department of Computing, Imperial College London, United Kingdom(计算系、帝国理工学院伦敦分校,英国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09182 2025-08-14 eess.IV cs.CV 79%

MedPatch: Confidence-Guided Multi-Stage Fusion for Multimodal Clinical Data

Baraa Al Jorf, Farah Shamout

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08925 2025-08-13 eess.AS cs.SD 79%

LPGNet: A Lightweight Network with Parallel Attention and Gated Fusion for Multimodal Emotion Recognition

Zhining He, Yang Xiao

机构 * Guangzhou University(广州大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 eess.AS

Comments Under peering review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01057 2025-08-13 cs.AI cs.RO 79%

Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance

Fengze Yang, Bo Yu, Yang Zhou, Xuewen Luo, Zhengzhong Tu, Chenxi Liu

机构 * Department of Civil & Environmental Engineering University of Utah(土木与环境工程系 犹他大学) Zachry Department of Civil and Environmental Engineering Texas A&M University(扎克里系 土木与环境工程系 德克萨斯农工大学) Department of Computer Science & Engineering Texas A&M University(计算机科学与工程系 德克萨斯农工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 24 pages, 6 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏