What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
视觉标记究竟编码了什么?揭示多模态大语言模型中的稀疏性与冗余性
机构 * Institute of Digital Twin, Eastern Institute of Technology(数字孪生研究所,东部技术研究所) ; Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) ; Shanghai Jiao Tong University(上海交通大学) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 文档图表理解 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本研究揭示多模态大语言模型中视觉标记的稀疏性与冗余性,通过EmbedLens工具发现活跃标记在进入模型前已编码细粒度信息,并提出中层注入方法提升效率。
Comments Accepted by CVPR2026