arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-19 至 2025-11-19 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 9 篇

2511.14604 2025-11-19 cs.CV 83%

XAttn-BMD: Multimodal Deep Learning with Cross-Attention for Femoral Neck Bone Mineral Density Estimation

Yilin Zhang, Leo D. Westbury, Elaine M. Dennison, Nicholas C. Harvey, Nicholas R. Fuggle, Rahman Attar

机构 * School of Electronics and Computer Science, University of Southampton, UK(电子与计算机科学学院,索姆塞特大学,英国) MRC Lifecourse Epidemiology Centre, University of Southampton, Southampton General Hospital, UK(生命课程流行病学研究中心,索姆塞特大学,南安普顿总医院,英国) NIHR Southampton Biomedical Research Centre, University of Southampton(南安普顿生物医学研究中心,索姆塞特大学) University Hospital NHS Foundation Trust, Southampton, UK(南安普顿国家健康服务基金会信托,英国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11 figures, 10 tables, 38 pages. Submitted to Artificial Intelligence in Medicine (currently with editor)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13755 2025-11-19 cs.LG cs.AI 83%

Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement

Zhe Yang, Wenrui Li, Hongtao Chen, Penghong Wang, Ruiqin Xiong, Xiaopeng Fan

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Zhengzhou Research Institute(哈尔滨工业大学郑州研究所) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究所) School of Mathematical Sciences, University of Electronic Science and Technology of China(数学学院,电子科学与技术大学) School of Electronic Engineering and Computer Science, Institute of Digital Media, Peking University(电子工程与计算机科学系,数字媒体研究所,北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13794 2025-11-19 cs.CV cs.AI 81%

FusionFM: All-in-One Multi-Modal Image Fusion with Flow Matching

Huayi Zhu, Xiu Shu, Youqiang Xiong, Qiao Liu, Rui Chen, Di Yuan, Xiaojun Chang, Zhenyu He

机构 * Guangzhou Institute of Technology, Xidian University(广州理工大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14693 2025-11-19 cs.CL 79%

Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances

Rishu Kumar Singh, Navneet Shreya, Sarmistha Das, Apoorva Singh, Sriparna Saha

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments To be published in the Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026 Special Track on AI for Social Impact )

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14157 2025-11-19 cs.CV 79%

Learning Representation and Synergy Invariances: A Povable Framework for Generalized Multimodal Face Anti-Spoofing

Xun Lin, Shuai Wang, Yi Yu, Zitong Yu, Jiale Zhou, Yizhong Liu, Xiaochun Cao, Alex Kot, Yefeng Zheng

机构 * Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学) Westlake University(西湖大学) Great Bay University(大亚湾大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15127 2025-11-19 cs.LG 79%

PRIMUS: Pretraining IMU Encoders with Multimodal Self-Supervision

Arnav M. Das, Chi Ian Tang, Fahim Kawsar, Mohammad Malekzadeh

机构 * Nokia Bell Labs Cambridge, UK(诺基亚贝尔实验室(剑桥,英国)) University of Washington, USA(华盛顿大学(美国)) University of Glasgow, UK(格拉斯哥大学(英国))

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Presented at ICASSP 2025. Also presented under the title "PRIMUS: Pretraining IMU Encoders with Multimodal and Self-Supervised Learning" at NeurIPS 2024 TSALM Workshop (Time Series in the Age of Large Models)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14601 2025-11-19 cs.CV cs.AI 62%

MRI Embeddings Complement Clinical Predictors for Cognitive Decline Modeling in Alzheimer's Disease Cohorts

Nathaniel Putera, Daniel Vilet Rodríguez, Noah Videcrantz, Julia Machnio, Mostafa Mehdipour Ghazi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at SPIE - Medical Imaging Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14698 2025-11-19 cs.CV cs.LG eess.SP 57%

HyMAD: A Hybrid Multi-Activity Detection Approach for Border Surveillance and Monitoring

Sriram Srinivasan, Srinivasan Aruchamy, Siva Ram Krisha Vadali

机构 * Sriram Srinivasan(独立研究者) Srinivasan Aruchamy(独立研究者) Siva Ram Krishna Vadali(独立研究者)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Multi-label seismic signal classification using novel attention-based feature fusion. Submitting to cs.CV due to relevance to general pattern recognition and time-frequency (spectrogram) analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14064 2025-11-19 cs.LG cs.AI stat.ME 57%

CafeMed: Causal Attention Fusion Enhanced Medication Recommendation

Kelin Ren, Chan-Yang Ju, Dong-Ho Lee

机构 * Department of Computer Science and Engineering, Hanyang University(韩阳大学计算机科学与工程系) Department of Applied Artificial Intelligence, Hanyang University(韩阳大学应用人工智能系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

Comments Accepted by BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏