arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-30 至 2025-10-30 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 8 篇

2510.24827 2025-10-30 cs.CV cs.MM 88%

MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition

Haoyang Zhang, Zhou Yang, Ke Sun, Yucai Pang, Guoliang Xu

机构 * Chongqing University of Posts and Telecommunications(重庆邮电大学) Xi’an Jiaotong University(西安交通大学) University of New South Wales(新南威尔士大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments The paper will be published in the MMAsia2025 conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09135 2025-10-30 cs.AI cs.CL cs.HC cs.LG 81%

Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation

Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K. Vail, Sunreeta Bhattacharya, Álvaro Fernández García, Kailana Baker-Matsuoka, Sheryl Mathew, Lori L. Holt, Fernando De la Torre

机构 * Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所) Department of Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校心理学系) Center for Perceptual Systems, The University of Texas at Austin(德克萨斯大学奥斯汀分校感知系统中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 22 pages, first three authors equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24777 2025-10-30 cs.CV cs.AI eess.IV 81%

Cross-Enhanced Multimodal Fusion of Eye-Tracking and Facial Features for Alzheimer's Disease Diagnosis

Yujie Nie, Jianzhang Ni, Yonglong Ye, Yuan-Ting Zhang, Yun Kwok Wing, Xiangqing Xu, Xin Ma, Lizhou Fan

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学) Engineering Research Center of Intelligent Unmanned System, Ministry of Education(智能无人机系统工程研究中心,教育部) Department of Psychiatry, The Chinese University of Hong Kong(心理学系,香港中文大学) Department of Electronic Engineering, The Chinese University of Hong Kong(电子工程系,香港中文大学) AICARE Lab, Guangdong Medical University(AICARE实验室,广东医科大学) Department of Neurology, Shandong University of Traditional Chinese Medicine Affiliated Hospital(神经内科,山东中医药大学附属医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 35 pages, 8 figures, and 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24919 2025-10-30 cs.CV cs.LG 79%

Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal Learning

Hossein R. Nowdeh, Jie Ji, Xiaolong Ma, Fatemeh Afghah

机构 * Holcombe Department of ECE(霍尔科姆电气与计算机工程系) Clemson University(克莱姆森大学) University of Arizona(亚利桑那大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03318 2025-10-30 cs.CV 79%

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning

Yibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang

机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Shanghai Innovation Institute(上海创新研究院) Shanghai AI Lab(上海人工智能实验室) Hunyuan, Tencent(腾讯 Hunyuan)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments [NeurIPS2025] Project Page: https://codegoat24.github.io/UnifiedReward/think

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24385 2025-10-30 cs.CV 57%

When are radiology reports useful for training medical image classifiers?

Herman Bergström, Zhongqi Yue, Fredrik D. Johansson

机构 * Department of Computer Science & Engineering, Chalmers University of Technology and University of Gothenburg(计算机科学与工程系,楚姆勒斯技术大学和哥德堡大学)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19638 2025-10-30 cs.CV 57%

HF-VTON: High-Fidelity Virtual Try-On via Consistent Geometric and Semantic Alignment

Ming Meng, Qi Dong, Jiajie Li, Zhe Zhu, Xingyu Wang, Zhaoxin Fan, Wei Zhao, Wenjun Wu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments After the publication of the paper, we discovered some significant errors/omissions that need to be corrected and improved

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11245 2025-10-30 cs.CV 57%

L2RSI: Cross-view LiDAR-based Place Recognition for Large-scale Urban Scenes via Remote Sensing Imagery

Ziwei Shi, Xiaoran Zhang, Wenjing Xu, Yan Xia, Yu Zang, Siqi Shen, Cheng Wang

机构 * Fujian Key Laboratory of Sensing and Computing for Smart Cities, Xiamen University, China(福建智能城市感知与计算重点实验室,厦门大学,中国) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China(多媒体可信感知与高效计算重点实验室,中华人民共和国教育部,厦门大学,中国) University of Science and Technology of China, China(中国科学技术大学,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 17 pages, 7 figures, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏