arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-04 至 2025-09-04 共收录 36 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 7 篇

2509.03433 2025-09-04 cs.CV 83%

Decoding Visual Neural Representations by Multimodal with Dynamic Balancing

Kaili sun, Xingyu Miao, Bing Zhai, Haoran Duan, Yang Long

机构 * Department of Computer Science, Durham University, UK.(计算机科学系,杜ham大学,英国) Computer and Information Sciences, Northumbria University, UK.(计算机与信息科学,北umbria大学,英国) Department of Automation, Tsinghua University, China.(自动化系,清华大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02943 2025-09-04 cs.IR 82%

Knowledge graph-based personalized multimodal recommendation fusion framework

Yu Fang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03408 2025-09-04 cs.CV cs.LG 79%

Scalable and Loosely-Coupled Multimodal Deep Learning for Breast Cancer Subtyping

Mohammed Amer, Mohamed A. Suliman, Tu Bui, Nuria Garcia, Serban Georgescu

机构 * Fujitsu Research of Europe Ltd(富士通欧洲研究有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03029 2025-09-04 cs.LG 78%

Multimodal learning of melt pool dynamics in laser powder bed fusion

Satyajit Mojumder, Pallock Halder, Tiana Tonge

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 20 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00905 2025-09-04 cs.CV cs.AI 62%

Spotlighter: Revisiting Prompt Tuning from a Representative Mining View

Yutong Gao, Maoyuan Shao, Xinyang Huang, Chuang Zhu, Lijuan Sun, Yu Weng, Xuan Liu, Guoshun Nan

机构 * School of Information Engineering, Minzu University of China(中国民族大学信息工程学院) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) National Library of China, Beijing, China(中国国家图书馆) School of Cyberspace Security, Beijing University of Posts and Telecommunications(北京邮电大学网络空间安全学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted as EMNLP 2025 Findings

Journal ref EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02602 2025-09-04 eess.IV 50%

Masked Autoencoder Pretraining and BiXLSTM ResNet Architecture for PET/CT Tumor Segmentation

Moona Mazher, Steven A Niederer, Abdul Qayyum

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏