arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2310.17373 2025-06-12 cs.IR 78%

Causality-Inspired Fair Representation Learning for Multimodal Recommendation

Weixin Chen, Li Chen, Yongxin Ni, Yuhan Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments In ACM Transactions on Information Systems (TOIS), 2025 (just accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07616 2025-06-10 cs.LG 78%

FuXi-Air: Urban Air Quality Forecasting Based on Emission-Meteorology-Pollutant multimodal Machine Learning

Zhixin Geng, Xu Fan, Xiqiao Lu, Yan Zhang, Guangyuan Yu, Cheng Huang, Qian Wang, Yuewu Li, Weichun Ma, Qi Yu, Libo Wu, Hao Li

机构 * Shanghai Key Laboratory of Atmospheric Particle Pollution and Prevention(上海大气颗粒污染与防治重点实验室) National Observations and Research Station for Wetland Ecosystems of the Yangtze Estuary(长江入海口湿地生态系统观测与研究站) Department of Environmental Science and Engineering(环境科学与工程学院) Shanghai Academy of AI for Science (SAIS)(上海人工智能科学研究院) MOE Laboratory for National Development and Intelligent Governance(教育部国家发展与智能治理实验室) Shanghai Institute for Energy and Carbon Neutrality Strategy(上海能源与碳中和战略研究院) IRDR ICoE on Risk Interconnectivity and Governance on Weather/Climate Extremes Impact and Public Health(国际减灾与发展研究院风险互联与治理国际联合实验室) Shanghai Institute of Eco-Chongming (SIEC)(上海生态崇明研究院) Shanghai Environmental Monitoring Center (SEMC)(上海市环境监测中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07532 2025-06-10 eess.SP 78%

A Unified Anti-Jamming Design in Complex Environments Based on Cross-Modal Fusion and Intelligent Decision-Making

Huake Wang, Xudong Han, Bairui Cai, Guisheng Liao, Yinghui Quan

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04870 2025-06-06 cs.LG 78%

Aligning Multimodal Representations through an Information Bottleneck

Antonio Almudévar, José Miguel Hernández-Lobato, Sameer Khurana, Ricard Marxer, Alfonso Ortega

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09392 2025-05-29 stat.ML cs.LG stat.CO 78%

Topological Eigenvalue Theorems for Tensor Analysis in Multi-Modal Data Fusion

Ronald Katende

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Error in Results. Need to re-run them

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12128 2025-05-28 cs.CE 78%

Multimodal Fusion with Relational Learning for Molecular Property Prediction

Zhengyang Zhou, Yunrui Li, Pengyu Hong, Hao Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14058 2025-05-27 cs.LG 78%

Towards the Causal Complete Cause of Multi-Modal Representation Learning

Jingyao Wang, Siyu Zhao, Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Fuchun Sun, Hui Xiong

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04452 2025-05-23 cs.IR 78%

COHESION: Composite Graph Convolutional Network with Dual-Stage Fusion for Multimodal Recommendation

Jinfeng Xu, Zheyu Chen, Wei Wang, Xiping Hu, Sang-Wook Kim, Edith C. H. Ngai

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11262 2025-05-19 cs.NE 78%

A Step towards Interpretable Multimodal AI Models with MultiFIX

Mafalda Malafaia, Thalea Schlender, Tanja Alderliesten, Peter A. N. Bosman

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 9 pages, 6 figures, submitted to GECCO conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10003 2025-05-16 cs.LG eess.SP 78%

AI2MMUM: AI-AI Oriented Multi-Modal Universal Model Leveraging Telecom Domain Large Model

Tianyu Jiao, Zhuoran Xiao, Yihang Huang, Chenhui Ye, Yijia Feng, Liyu Cai, Jiang Chang, Fangkun Liu, Yin Xu, Dazhi He, Yunfeng Guan, Wenjun Zhang

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学 cooperative medianet innovation center) Nokia Bell Labs(诺基亚贝尔实验室) Institute of Intelligent Communications and Network Security, Chongqing University of Posts and Telecommunications(重庆邮电大学智能通信与网络安全研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18024 2025-05-16 cs.LG 78%

Multimodal Learning with Uncertainty Quantification based on Discounted Belief Fusion

Grigor Bezirganyan, Sana Sellami, Laure Berti-Équille, Sébastien Fournier

机构 * Aix-Marseille Univ(艾克斯-马赛大学) LIS(实验室) IRD(法国国家科研 Institute) ESPACE-DEV

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Journal ref Proceedings of The 28th International Conference on Artificial Intelligence and Statistics 2025, in Proceedings of Machine Learning Research 258:3142-3150 Available from https://proceedings.mlr.press/v258/bezirganyan25a.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04120 2025-05-14 cs.LG 78%

Transformer representation learning is necessary for dynamic multi-modal physiological data on small-cohort patients

Bingxu Wang, Min Ge, Kunzhi Cai, Yuqi Zhang, Zeyi Zhou, Wenjiao Li, Yachong Guo, Wei Wang, Qing Zhou

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05698 2025-05-12 cs.LG 78%

Unsupervised Multi-modal Feature Alignment for Time Series Representation Learning

Chen Liang, Donghua Yang, Zhiyu Liang, Hongzhi Wang, Zheng Liang, Xiyang Zhang, Jianfeng Huang

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04634 2025-05-09 cs.LG cs.CE 78%

MatMMFuse: Multi-Modal Fusion model for Material Property Prediction

Abhiroop Bhattacharya, Sylvain G. Cloutier

机构 * Department of Electrical Engineering, École de technologie supérieure, Montréal, Canada(电气工程系,超技术学院,蒙特利尔,加拿大)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Presented at AI for Accelerated Materials Design(AI4Mat), ICLR 2025 (https://openreview.net/forum?id=pN4Zg6HBlq#discussion)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02168 2025-05-06 cs.AR 78%

CircuitFusion: Multimodal Circuit Representation Learning for Agile Chip Design

Wenji Fang, Shang Liu, Jing Wang, Zhiyao Xie

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by ICLR 2025 (https://openreview.net/forum?id=rbnf7oe6JQ)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01945 2025-05-06 cs.MA cs.RO 78%

Act Natural! Extending Naturalistic Projection to Multimodal Behavior Scenarios

Hamzah I. Khan, David Fridovich-Keil

机构 * Department of Aerospace Engineering and Engineering Mechanics , University of Texas at Austin(航空航天工程与工程力学系,德克萨斯大学奥斯汀分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01135 2025-05-05 cs.LG 78%

Dual-Forecaster: A Multimodal Time Series Model Integrating Descriptive and Predictive Texts

Wenfa Wu, Guanyu Zhang, Zheng Tan, Yi Wang, Hongsheng Qi

机构 * Lenovo Research(联想研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00842 2025-05-05 cs.RO cs.SY eess.SY math.GR 78%

Fault-Tolerant Multi-Modal Localization of Multi-Robots on Matrix Lie Groups

Mahboubeh Zarei, Robin Chhabra

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00176 2025-05-02 cs.CE 78%

Generative Multimodal Multiscale Data Fusion for Digital Twins in Aerosol Jet Electronics Printing

Fatemeh Elhambakhsh, Suk Ki Lee, Hyunwoong Ko

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21826 2025-05-01 cs.RO 78%

An Underwater, Fault-Tolerant, Laser-Aided Robotic Multi-Modal Dense SLAM System for Continuous Underwater In-Situ Observation

Yaming Ou, Junfeng Fan, Chao Zhou, Pengju Zhang, Zongyuan Shen, Yichen Fu, Xiaoyan Liu, Zengguang Hou

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16524 2025-04-24 cs.IR 78%

Modality Reliability Guided Multimodal Recommendation

Xue Dong, Xuemeng Song, Na Zheng, Sicheng Zhao, Guiguang Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14927 2025-04-22 cs.HC 78%

Multimodal Non-Semantic Feature Fusion for Predicting Segment Access Frequency in Lecture Archives

Ruozhu Sheng, Jinghong Li, Shinobu Hasegawa

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 16 pages, 7 figures. Preliminary version; work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14833 2025-04-22 cs.NI cs.CR 78%

IoT-AMLHP: Aligned Multimodal Learning of Header-Payload Representations for Resource-Efficient Malicious IoT Traffic Classification

Fengyuan Nie, Guangjie Liu, Weiwei Liu, Jianan Huang, Bo Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13465 2025-04-21 cs.LG 78%

Are you SURE? Enhancing Multimodal Pretraining with Missing Modalities through Uncertainty Estimation

Duy A. Nguyen, Quan Huu Do, Khoa D. Doan, Minh N. Do

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11490 2025-04-18 cs.LG stat.ME 78%

Interventional Imbalanced Multi-Modal Representation Learning via $β$-Generalization Front-Door Criterion

Yi Li, Fei Song, Changwen Zheng, Jiangmeng Li, Fuchun Sun, Hui Xiong

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12025 2025-04-17 cs.LG 78%

FedEPA: Enhancing Personalization and Modality Alignment in Multimodal Federated Learning

Yu Zhang, Qingfeng Du, Jiaqi Lv

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09260 2025-04-15 cs.AR cs.LG 78%

NetTAG: A Multimodal RTL-and-Layout-Aligned Netlist Foundation Model via Text-Attributed Graph

Wenji Fang, Wenkai Li, Shang Liu, Yao Lu, Hongce Zhang, Zhiyao Xie

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted by Design Automation Conference (DAC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12501 2025-04-03 cs.IR 78%

Improving Multi-modal Recommender Systems by Denoising and Aligning Multi-modal Content and User Feedback

Guipeng Xv, Xinyu Li, Ruobing Xie, Chen Lin, Chong Liu, Feng Xia, Zhanhui Kang, Leyu Lin

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments After further review, we believe the content of the paper is not yet fully ready and requires additional time for improvement. To ensure quality, we have decided to withdraw this preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21964 2025-03-31 cs.LG q-bio.NC 78%

NeuroLIP: Interpretable and Fair Cross-Modal Alignment of fMRI and Phenotypic Text

Yanting Yang, Xiaoxiao Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20823 2025-03-28 cs.CR 78%

Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy

Joonhyun Jeong, Seyun Bae, Yeonsung Jung, Jaeryong Hwang, Eunho Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏