arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-10 至 2025-11-10 共收录 3 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 3 篇

2511.05474 2025-11-10 cs.CV 79%

Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection

Xian-Hong Huang, Hui-Kai Su, Chi-Chia Sun, Jun-Wei Hsieh

机构 * Department of Electrical Engineering, National Formosa University, Taiwan(台湾国立Formosa大学电子工程系) Department of Electrical Engineering, National Taipei University, Taiwan(台湾国立台北大学电子工程系) College of Artificial Intelligence and Green Energy, National Yang Ming Chiao Tung University, Taiwan(台湾国立阳明交通大学人工智能与再生能源学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05044 2025-11-10 cs.CV 57%

Medical Referring Image Segmentation via Next-Token Mask Prediction

Xinyu Chen, Yiran Wang, Gaoyang Pang, Jiafu Hao, Chentao Yue, Luping Zhou, Yonghui Li

机构 * School of Electrical and Computer Engineering, University of Sydney(悉尼大学电气与计算机工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE Transactions on Medical Imaging for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04790 2025-11-10 cs.LG cs.AI stat.ML 57%

Causal Structure and Representation Learning with Biomedical Applications

Caroline Uhler, Jiaqi Zhang

机构 * Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology(电气工程与计算机科学系,麻省理工学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments This article has successfully completed peer review and will appear in the Proceedings of the International Congress of Mathematicians 2026. Both authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏