arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-10-23 至 2025-10-23 共收录 5 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 5 篇

2510.19160 2025-10-23 cs.LG 83%

Preliminary Use of Vision Language Model Driven Extraction of Mouse Behavior Towards Understanding Fear Expression

Paimon Goulart, Jordan Steinhauser, Kylene Shuler, Edward Korzus, Jia Chen, Evangelos E. Papalexakis

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 VLM训练与架构 :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01822 2025-10-23 cs.CV 83%

VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models

Byung-Kwan Lee, Ryo Hachiuma, Yu-Chiang Frank Wang, Yong Man Ro, Yueh-Hua Wu

机构 * NVIDIA KAIST(韩国科学技术院) National Taiwan University(国立台湾大学)

专题命中 VLM训练与架构 :vision language model(title);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments CVPR 2025, Project page: https://byungkwanlee.github.io/VLsI-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19802 2025-10-23 cs.CV 79%

Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models

Xiaozhen Qiao, Jingkai Zhao, Yuqiu Jiang, Xianda Guo, Zhe Sun, Hongyuan Zhang, Xuelong Li

机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Institute of Artificial Intelligence (TeleAI), China Telecom, P. R. China(人工智能研究所(TeleAI),中国电信,中华人民共和国) College of Computer Science, Wuhan University(计算机科学学院,武汉大学) School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院,光学与电子学(iOPEN),西北工业大学)

专题命中 VLM训练与架构 :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19678 2025-10-23 cs.CV cs.AI 62%

I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs

John Burden, Jonathan Prunty, Ben Slater, Matthieu Tehenan, Greg Davis, Lucy Cheke

机构 * Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能研究中心、剑桥大学) Department of Engineering, University of Cambridge(工程系、剑桥大学) Department of Psychology, University of Cambridge(心理学系、剑桥大学) Department of Computer Science, University of Cambridge(计算机科学系、剑桥大学)

专题命中 VLM训练与架构 :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17452 2025-10-23 cs.CV cs.AI 62%

Training-Free Label Space Alignment for Universal Domain Adaptation

Dujin Lee, Sojung An, Jungmyung Wi, Kuniaki Saito, Donghyun Kim

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments 22 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏