arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-31 至 2025-10-31 共收录 32 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 4 篇

2506.05696 2025-10-31 cs.CV 70%

MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory

Ana Carolina Condez, Diogo Tavares, João Magalhães

机构 * NOVA LINCS, NOVA School of Science and Technology(NOVA LINCS,NOVA科学与技术学院)

专题命中 VLM训练与架构 :vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Updated version: corresponds to the ACM MM '25 published paper and includes full appendix material

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01074 2025-10-31 cs.LG 57%

Omni-Mol: Multitask Molecular Model for Any-to-any Modalities

Chengxin Hu, Hao Li, Yihe Yuan, Zezheng Song, Chenyang Zhao, Haixin Wang

机构 * National University of Singapore(新加坡国立大学) University of Maryland, College Park(马里兰大学 College Park 分校) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.LG

Comments 44 pages, 9 figures, 13 tables, paper accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏