ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) ; University of Chinese Academy of Sciences(中国科学院大学)
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.CV
Comments Published as a conference paper at ICCV 2025. Project page: https://github.com/icip-cas/ShortV