arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-09-25 至 2025-09-25 共收录 36 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 5 篇

2502.18485 2025-09-25 q-bio.NC cs.CV 85%

Deciphering Functions of Neurons in Vision-Language Models

Jiaqi Xu, Cuiling Lan, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV

Comments Accepted by the 31st ACM International Conference on Multimedia (ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13739 2025-09-25 cs.CV 79%

Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector

Yiming Cao, Yanjie Li, Kaisheng Liang, Bin Xiao

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19524 2025-09-25 cs.AI cs.RO 79%

Score the Steps, Not Just the Goal: VLM-Based Subgoal Evaluation for Robotic Manipulation

Ramy ElMallah, Krish Chhajer, Chi-Guhn Lee

机构 * University of Toronto(多伦多大学)

专题命中 VLM训练与架构 :VLM(title);vision-language model(abstract);分类 cs.AI

Comments Accepted to the CoRL 2025 Eval&Deploy Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21976 2025-09-25 cs.AI 70%

Compression Strategies for Efficient Multimodal LLMs in Medical Contexts

Tanvir A. Khan, Aranya Saha, Ismam N. Swapnil, Mohammad A. Haque

机构 * Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology (BUET)(电子与电气工程系,孟加拉国工程与技术大学)

专题命中 VLM训练与架构 :LLaVA(abstract);multimodal large language model(abstract);分类 cs.AI

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19924 2025-09-25 cs.LG cs.AI 62%

Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches

Remo Sasso, Michelangelo Conserva, Dominik Jeurissen, Paulo Rauber

机构 * School of Electronic Engineering and Computer Science(电子工程与计算机科学学院)

专题命中 VLM训练与架构 :VLM(abstract);分类 cs.AI、cs.LG

Comments 16 pages, 7 figures. Accepted for presentation at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on the Foundations of Reasoning in Language Models (FoRLM)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他VLM 1 篇

2509.20279 2025-09-25 cs.CV q-bio.QM 57%

A co-evolving agentic AI system for medical imaging analysis

Songhao Li, Jonathan Xu, Tiancheng Bao, Yuxuan Liu, Yuchen Liu, Yihang Liu, Lilin Wang, Wenhui Lei, Sheng Wang, Yinuo Xu, Yan Cui, Jialu Yao, Shunsuke Koga, Zhi Huang

机构 * Department of Pathology and Laboratory Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学) Department of Electrical and System Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学) The Wharton School, University of Pennsylvania(沃顿商学院,宾夕法尼亚大学) Department of Bioengineering, University of Pennsylvania(生物工程系,宾夕法尼亚大学) Department of Computer and Information Science, University of Pennsylvania(计算机与信息科学系,宾夕法尼亚大学) Department of Biostatistics, Epidemiology & Informatics, University of Pennsylvania(生物统计学、流行病学与信息学系,宾夕法尼亚大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏