arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-04 至 2025-11-04 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2511.01444 2025-11-04 cs.AI 83%

Robust Multimodal Sentiment Analysis via Double Information Bottleneck

Huiting Huang, Tieliang Gong, Kai He, Jialun Wu, Erik Cambria, Mengling Feng

机构 * School of Computer Science(计算机科学学院) Technology, Xi’an Jiaotong University(技术,西安交通大学) Shaanxi Provincial Key Laboratory of Big Data Knowledge Engineering, Xi’an Jiaotong University(大数据知识工程省级重点实验室,西安交通大学) Saw Swee Hock School of Public Health, National University of Singapore(Saw Swee Hock 公共卫生学院,新加坡国立大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) School of Computer Science, Northwestern Polytechnical University(计算机科学学院,西北工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01320 2025-11-04 cs.AI 83%

OmniFuser: Adaptive Multimodal Fusion for Service-Oriented Predictive Maintenance

Ziqi Wang, Hailiang Zhao, Yuhao Yang, Daojiang Hu, Cheng Bao, Mingyi Liu, Kai Di, Schahram Dustdar, Zhongjie Wang, Shuiguang Deng

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Hangzhou School of Automation, Zhejiang Normal University(浙江师范大学杭州自动化学院) Distributed Systems Group at the TU Wien and with ICREA at the UPF, Barcelona(维也纳大学分布式系统组和巴塞罗那大学ICREA)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07841 2025-11-04 cs.NI cs.LG 82%

Task-Oriented Multimodal Token Transmission in Resource-Constrained Multiuser Networks

Junhe Zhang, Wanli Ni, Pengwei Wang, Dongyu Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00987 2025-11-04 cs.LG 82%

Balanced Multimodal Learning via Mutual Information

Rongrong Xie, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati (SISSA)(国际先进研究学院(SISSA))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01435 2025-11-04 cs.CV 79%

Contrast-Guided Cross-Modal Distillation for Thermal Object Detection

SiWoo Kim, JhongHyun An

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00859 2025-11-04 cs.CV 79%

Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion

Jaehyun Park, Konyul Park, Daehun Kim, Junseo Park, Jun Won Choi

机构 * Seoul National University(首尔国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19769 2025-11-04 cs.CV 79%

AIM: Adaptive Intra-Network Modulation for Balanced Multimodal Learning

Shu Shen, C. L. Philip Chen, Tong Zhang

机构 * Guangdong Provincial Key Laboratory of Computational AI Models and Cognitive Intelligence(广东省计算人工智能模型与认知智能重点实验室) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) Pazhou Lab(琶洲实验室) Engineering Research Center of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human(教育部健康智能感知与并行数字人工程研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 13pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01357 2025-11-04 cs.CV cs.AI 79%

CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering

Qiangguo Jin, Xianyao Zheng, Hui Cui, Changming Sun, Yuqi Fang, Cong Cong, Ran Su, Leyi Wei, Ping Xuan, Junbo Wang

机构 * School of Software, Northwestern Polytechnical University, Shaanxi, China(西北工业大学软件学院) Yangtze River Delta Research Institute of Northwestern Polytechnical University, Taicang, China(西北工业大学长江三角研究 institute) Department of Computer Science and Information Technology, La Trobe University, Melbourne, Australia(拉筹伯大学计算机科学与信息技术系) CSIRO Data61, Sydney, Australia(CSIRO Data61) School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院) Australian Institute of Health Innovation (AIHI), Macquarie University, Australia(麦考瑞大学健康创新研究所) School of Computer Software, College of Intelligence and Computing, Tianjin University, Tianjin, China(天津大学计算机软件学院) Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University, Macao Special Administrative Region of China(澳门理工学院人工智能驱动药物发现中心) Department of Computer Science, School of Engineering, Shantou University, Guangdong, China(汕头大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments The paper has been accepted by the 33rd Pacific Conference on Computer Graphics and Applications (Pacific Graphics 2025)

Journal ref PG2025 Conference Papers, Posters, and Demos, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00949 2025-11-04 cs.LG 78%

Motion-Robust Multimodal Fusion of PPG and Accelerometer Signals for Three-Class Heart Rhythm Classification

Yangyang Zhao, Matti Kaisti, Olli Lahdenoja, Tero Koivisto

机构 * Department of Computing, Faculty of Technology, University of Turku(图波大学计算系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted for publication in the Companion of the 2025 ACM International Joint Conference on Pervasive and Ubiquitous Computing and the 2025 International Symposium on Wearable Computers (UbiComp/ISWC 2025 Companion). 5 pages, 3 figures. Author's accepted manuscript (AAM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00509 2025-11-04 cs.AI cs.CR 77%

Reimagining Safety Alignment with An Image

Yifan Xia, Guorui Chen, Wenqian Yu, Zhijiang Li, Philip Torr, Jindong Gu

机构 * School of Information Management, Wuhan University, Wuhan, China(武汉大学信息管理学院) Torr Vision Group, University of Oxford, Oxford, United Kingdom(牛津大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01463 2025-11-04 cs.CV cs.AI cs.GR 73%

HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA

Lei Hu, Yongjing Ye, Shihong Xia

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5figures. The Thirty-Ninth Annual Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01284 2025-11-04 cs.CV cs.AI 73%

Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions

Karma Phuntsho, Abdullah, Kyungmi Lee, Ickjai Lee, Euijoon Ahn

机构 * James Cook University(詹姆斯库克大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01082 2025-11-04 cs.CV cs.AI cs.LG 73%

GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction

Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Cyrus Shahabi

机构 * of Computer Science, University of Southern California, Los Angeles, CA, USA Computer Engineering, University of Southern California, Los Angeles, CA, USA

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to IEEE International Conference on Data Mining (ICDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00218 2025-11-04 cs.CV cs.AI 62%

DM-QPMNET: Dual-modality fusion network for cell segmentation in quantitative phase microscopy

Rajatsubhra Chakraborty, Ana Espinosa-Momox, Riley Haskin, Depeng Xu, Rosario Porras-Aguilar

机构 * College of Computing and Informatics, University of North Carolina at Charlotte, NC, USA(计算与信息学院,北卡罗来纳大学夏洛特分校) Department of Physics and Optical Science, University of North Carolina at Charlotte, NC, USA(物理与光学科学系,北卡罗来纳大学夏洛特分校) Center for TAIMing AI, University of North Carolina at Charlotte, NC, USA(TAIMing AI中心,北卡罗来纳大学夏洛特分校)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26466 2025-11-04 cs.CV cs.LG 57%

Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition

Pei Peng, MingKun Xie, Hang Hao, Tong Jin, ShengJun Huang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19739 2025-11-04 cs.CV 57%

FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion

Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen, Wenbo Zhao, Kui Jiang, Xianming Liu

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Zhengzhou Research Institute, Harbin Institute of Technology(郑州研究院,哈尔滨工业大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments [IROS 2025, camera ready version]: 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏