arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2509.03837 2025-09-05 cs.LG cs.IT math.IT 78%

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

Kimia Ehsani, Walid Saad

机构 * Bradley Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at IEEE GLOBECOM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03029 2025-09-04 cs.LG 78%

Multimodal learning of melt pool dynamics in laser powder bed fusion

Satyajit Mojumder, Pallock Halder, Tiana Tonge

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 20 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02275 2025-09-03 cs.RO 78%

Human-Inspired Soft Anthropomorphic Hand System for Neuromorphic Object and Pose Recognition Using Multimodal Signals

Fengyi Wang, Xiangyu Fu, Nitish Thakor, Gordon Cheng

机构 * Institute for Cognitive Systems, Technical University of Munich(认知系统研究所,慕尼黑技术大学) Department of Biomedical Engineering, Johns Hopkins University(生物医学工程系,约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18551 2025-08-27 cs.LG 78%

BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integration

Jun Hou, Le Wang, Xuan Wang

机构 * Department of Computer Science, Virginia Tech, Blacksburg, VA, USA(计算机科学系,弗吉尼亚理工学院,布莱克斯堡,VA,美国) Department of Agricultural and Applied Economics, Virginia Tech, Blacksburg, VA, USA(农业与应用经济学系,弗吉尼亚理工学院,布莱克斯堡,VA,美国)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17640 2025-08-26 eess.SP 78%

Multimodal Radio and Vision Fusion for Robust Localization in Urban V2I Communications

Can Zheng, Jiguang He, Chung G. Kang, Guofa Cai, Henk Wymeersch

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 6 pages, 6 figures, submitted to conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14485 2025-08-22 cs.IR 78%

Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion

Moyu Zhang, Yongxiang Tang, Yujun Jin, Jinxin Hu, Yu Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by CIKM 2025, 11 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14844 2025-08-21 cs.LG 78%

Multimodal Quantum Vision Transformer for Enzyme Commission Classification from Biochemical Representations

Murat Isik, Mandeep Kaur Saggi, Humaira Gowher, Sabre Kais

机构 * Purdue University(普渡大学) NC State University(北卡罗来纳州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at IEEE International Conference on Quantum Artificial Intelligence (QAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11762 2025-08-21 cs.LG cs.RO 78%

MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations

Daniel Bogdoll, Yitian Yang, Tim Joseph, Melih Yazgan, J. Marius Zöllner

机构 * FZI Research Center for Information Technology, Germany(德国弗赖堡信息科技研究中心) Karlsruhe Institute of Technology, Germany(德国卡尔斯鲁厄理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Daniel Bogdoll and Yitian Yang contributed equally. Accepted for publication at IV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10753 2025-08-15 cs.IR 78%

Hypercomplex Prompt-aware Multimodal Recommendation

Zheyu Chen, Jinfeng Xu, Hewei Wang, Shuo Yang, Zitong Wan, Haibo Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10696 2025-08-15 cs.CE 78%

Chem3DLLM: 3D Multimodal Large Language Models for Chemistry

Lei Jiang, Shuzhou Sun, Biqing Qi, Yuchen Fu, Xiaohua Xu, Yuqiang Li, Dongzhan Zhou, Tianfan Fu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09664 2025-08-14 cs.IR 78%

Multimodal Fusion And Sparse Attention-based Alignment Model for Long Sequential Recommendation

Yongrui Fu, Jian Liu, Tao Li, Zonggang Wu, Shouke Qin, Hanmeng Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15250 2025-08-13 cs.LG 78%

Multi-modal Policies with Physics-informed Representations in Complex Fluid Environments

Haodong Feng, Peiyan Hu, Yue Wang, Dixia Fan

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08257 2025-08-13 eess.SP cs.RO 78%

Where is the Boundary: Multimodal Sensor Fusion Test Bench for Tissue Boundary Delineation

Zacharias Chen, Alexa Cristelle Cahilig, Sarah Dias, Prithu Kolar, Ravi Prakash, Patrick J. Codd

机构 * Duke University(杜克大学) Research Triangle High School(研究三角高中) North Carolina School of Science and Mathematics(北卡罗来纳科学与数学高中)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 4 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07536 2025-08-12 cs.LG 78%

Physics-Informed Multimodal Bearing Fault Classification under Variable Operating Conditions using Transfer Learning

Tasfiq E. Alam, Md Manjurul Ahsan, Shivakumar Raman

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.06367 2025-08-08 cs.LG 78%

Does a Technique for Building Multimodal Representation Matter? -- Comparative Analysis

Maciej Pawłowski, Anna Wróblewska, Sylwia Sysko-Romańczuk

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22880 2025-07-31 cs.IR 78%

AUV-Fusion: Cross-Modal Adversarial Fusion of User Interactions and Visual Perturbations Against VARS

Hai Ling, Tianchi Wang, Xiaohao Liu, Zhulin Tao, Lifang Yang, Xianglin Huang

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

Comments 14 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22426 2025-07-31 cs.LG 78%

Multimodal Late Fusion Model for Problem-Solving Strategy Classification in a Machine Learning Game

Clemens Witt, Thiemo Leonhardt, Nadine Bergner, Mareen Grillenberger

机构 * TUD Dresden University of Technology(德累斯顿技术大学) RWTH Aachen University(亚琛工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments This is the author's version of a paper accepted for publication at the 2025 European Conference on Technology Enhanced Learning (EC-TEL 2025). The final authenticated version will be published in the Lecture Notes in Computer Science (LNCS) series by Springer and will be available via SpringerLink

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10091 2025-07-30 eess.IV 78%

G$^{2}$SF-MIAD: Geometry-Guided Score Fusion for Multimodal Industrial Anomaly Detection

Chengyu Tao, Xuanming Cao, Juan Du

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16088 2025-07-25 astro-ph.IM astro-ph.HE 78%

Applying multimodal learning to Classify transient Detections Early (AppleCiDEr) I: Data set, methods, and infrastructure

Alexandra Junell, Argyro Sasli, Felipe Fontinele Nunes, Maojie Xu, Benny Border, Nabeel Rehemtulla, Mariia Rizhko, Yu-Jing Qin, Theophile Jegou Du Laz, Antoine Le Calloch, Sushant Sharma Chaudhary, Shaowei Wu, Jesper Sollerman, Niharika Sravan, Steven L. Groom, David Hale, Mansi M. Kasliwal, Josiah Purdum, Avery Wold, Matthew J. Graham, Michael W. Coughlin

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15769 2025-07-22 cs.LG 78%

Multi-Modal Sensor Fusion for Proactive Blockage Prediction in mmWave Vehicular Networks

Ahmad M. Nazar, Abdulkadir Celik, Mohamed Y. Selim, Asmaa Abdallah, Daji Qiao, Ahmed M. Eltawil

机构 * Department of Electrical and Computer Engineering, Iowa State University of Science and Technology(电气与计算机工程系,爱荷华州立大学) Computer, Electrical, and Mathematical Sciences & Engineering (CEMSE) Division at King Abdullah University of Science and Technology (KAUST)(国王 Abdullah 科学与技术大学(KAUST)计算机、电气和数学科学与工程(CEMSE)系)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Accepted in IEEE Asilomar Conference on Signals, Systems, and Computers 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14361 2025-07-22 cs.IR 78%

RaMen: Multi-Strategy Multi-Modal Learning for Bundle Construction

Huy-Son Nguyen, Quang-Huy Nguyen, Duc-Hoang Pham, Duc-Trong Le, Hoang-Quynh Le, Padipat Sitkrongwong, Atsuhiro Takasu, Masoud Mansoury

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11870 2025-07-17 cs.CE 78%

MNO : A Multi-modal Neural Operator for Parametric Nonlinear BVPs

Vamshi C. Madala, Nithin Govindarajan, Shivkumar Chandrasekaran

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09998 2025-07-15 cs.IR 78%

SLIF-MR: Self-loop Iterative Fusion of Heterogeneous Auxiliary Information for Multimodal Recommendation

Jie Guo, Jiahao Jiang, Ziyuan Guo, Bin Song, Yue Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 10 pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07367 2025-07-11 q-bio.BM cs.LG 78%

Platform for Representation and Integration of multimodal Molecular Embeddings

Erika Yilin Zheng, Yu Yan, Baradwaj Simha Sankar, Ethan Ji, Steven Swee, Irsyad Adam, Ding Wang, Alexander Russell Pelletier, Alex Bui, Wei Wang, Peipei Ping

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19960 2025-07-10 cs.LG 78%

SeisMoLLM: Advancing Seismic Monitoring via Cross-modal Transfer with Pre-trained Large Language Model

Xinghao Wang, Feng Liu, Rui Su, Zhihui Wang, Lihua Fang, Lianqing Zhou, Lei Bai, Wanli Ouyang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) International School of Information Science and Engineering(国际信息科学与工程学院) School of Electronic Information and Electrical Engineering(电子信息与电气工程学院) Shanghai Jiao Tong University(上海交通大学) Dalian University of Technology(大连理工大学) Institute of Earthquake Forecasting(地震预测研究所) China Earthquake Administration(中国地震局)

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

Comments Code is available at https://github.com/StarMoonWang/SeisMoLLM. v2 fixed errors in the location figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01337 2025-07-03 cs.IT eess.SP math.IT 78%

Dynamical Multimodal Fusion with Mixture-of-Experts for Localizations

Bohao Wang, Zitao Shuai, Fenghao Zhu, Chongwen Huang, Yongliang Shen, Zhaoyang Zhang, Qianqian Yang, Sami Muhaidat, Merouane Debbah

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22364 2025-06-30 cs.RO 78%

Robotic Multimodal Data Acquisition for In-Field Deep Learning Estimation of Cover Crop Biomass

Joe Johnson, Phanender Chalasani, Arnav Shah, Ram L. Ray, Muthukumar Bagavathiannan

机构 * Department of Soil and Crop Sciences, Texas A&M University(土壤与作物科学系,德克萨斯A&M大学) Department of Computer Science and Engineering, Texas A&M University(计算机科学与工程系,德克萨斯A&M大学) Department of Mechanical Engineering, Texas A&M University(机械工程系,德克萨斯A&M大学) College of Agriculture, Food and Natural Resources, Prairie View A&M University(农业、食品与自然资源学院,普里奥里视A&M大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted in the Extended Abstract, The 22nd International Conference on Ubiquitous Robots (UR 2025), Texas, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19077 2025-06-25 cs.RO 78%

Multimodal Anomaly Detection with a Mixture-of-Experts

Christoph Willibald, Daniel Sliwowski, Dongheui Lee

机构 * Institute of Robotics and Mechatronics (DLR), German Aerospace Center(机器人与机电研究所(DLR),德国航空航天中心) Autonomous Systems Lab, Institute of Computer Technology, TU Wien(自主系统实验室,计算机技术研究所,维也纳技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 8 pages, 5 figures, 1 table, the paper has been accepted for publication in the Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17068 2025-06-23 q-bio.NC cs.ET eess.SP 78%

Cross-Modal Epileptic Signal Harmonization: Frequency Domain Mapping Quantization for Pre-training a Unified Neurophysiological Transformer

Runkai Zhang, Hua Yu, John Q. Gan, Haixian Wang

专题命中 多模态训练与对齐 :cross-modal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16968 2025-06-23 cs.CR cs.CY 78%

MM-AttacKG: A Multimodal Approach to Attack Graph Construction with Large Language Models

Yongheng Zhang, Xinyun Zhao, Yunshan Ma, Haokai Ma, Yingxiao Guan, Guozheng Yang, Yuliang Lu, Xiang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏