arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-12 至 2025-09-12 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 9 篇

2505.19455 2025-09-12 cs.CV cs.AI cs.LG 84%

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

Xu Li, Fan Lyu

机构 * Khoury College of Computer Sciences, Northeastern University(东北大学克劳尔计算机科学学院) New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09427 2025-09-12 cs.CV 83%

FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

Yuchan Jie, Yushen Xu, Xiaosong Li, Fuqiang Zhou, Jianming Lv, Huafeng Li

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Physics and Optoelectronic Engineering, Foshan University(佛山大学物理与光电工程学院) School of Instrumentation Science and Optelectronics Engineering, Beihang University(北航仪器科学与光电工程学院) School of Information Engineering and Automation, Kunming University of Science and Technology(昆明理工大学信息工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref Information Fusion, 2025, 121: 103146

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09114 2025-09-12 cs.IR 82%

Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation

Kelin Ren, Chan-Yang Ju, Dong-Ho Lee

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09290 2025-09-12 cs.CV cs.AI 82%

Modality-Agnostic Input Channels Enable Segmentation of Brain lesions in Multimodal MRI with Sequences Unavailable During Training

Anthony P. Addison, Felix Wagner, Wentian Xu, Natalie Voets, Konstantinos Kamnitsas

机构 * Department of Engineering Science, University of Oxford, Oxford, UK(工程科学系,牛津大学,牛津,英国) Nuffield Department of Clinical Neurosciences, University of Oxford, Oxford, UK(牛津大学临床神经科学系,牛津,英国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to MICCAI 2025, for the following workshop: ML-CDS 2025: Multimodal Learning and Fusion Across Scales for Clinical Decision Support

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09064 2025-09-12 cs.CV 79%

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

Qiuhui Chen, Xuancheng Yao, Huping Ye, Yi Hong

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Journal of Biomedical and Health Informatics (JBHI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14707 2025-09-12 cs.CY 71%

Information Fusion in Multimodal IoT Systems for physical activity level monitoring

Mohsen Shirali, Zahra Ahmadi, Jose-Luis Bayo-Monton, Zoe Valero-Ramon, Carlos Fernandez-Llatas

专题命中 多模态训练与对齐 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06465 2025-09-12 cs.LG cs.CE q-bio.BM 67%

CAME-AB: Cross-Modality Attention with Mixture-of-Experts for Antibody Binding Site Prediction

Hongzong Li, Jiahao Ma, Zhanpeng Shi, Rui Xiao, Fanming Jin, Ye-Fan Hu, Hangjun Che, Jian-Dong Huang

机构 * Generative AI Research and Development Center The Hong Kong University of Science and Technology(生成式人工智能研究与发展中心 香港科学大学) MILES The University of Hong Kong(MILES 香港大学) College of Veterinary Medicine Jilin University(吉林大学兽医学院) School of Chemistry and Chemical Engineering South China University of Technology(华南理工大学化学与化工学院) School of Biomedical Sciences The University of Hong Kong(香港大学生物医学科学学院) Computational Immunology Centre BayVax Biotech Limited(计算免疫学中心 BayVax 生物技术有限公司) College of Electronic and Information Engineering Southwest University(西南大学电子与信息工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05211 2025-09-12 cs.CV 57%

VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization

Sihan Yang, Runsen Xu, Chenhang Cui, Tai Wang, Dahua Lin, Jiangmiao Pang

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08830 2025-09-12 eess.SP cs.LG 50%

A Masked Representation Learning to Model Cardiac Functions Using Multiple Physiological Signals

Seong-A Park, Jong-Eui Chae, Sungdong Kim, Hyung-Chul Lee, Hyun-Lim Yang

机构 * Seoul National University Hospital(首尔国立大学医院) Seoul National University(首尔国立大学) KAIST AI(韩国科学技术院人工智能研究所)

专题命中 多模态训练与对齐 :multi-modal(abstract)

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏