arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAViE:用于细粒度视觉感知和高效多模态推理的多尺度自适应视觉编码器

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

Shaofei Lei

arXiv 2607.24424首次发表:更新:

AI 中文总结

研究针对视觉语言模型问题,提出多尺度自适应视觉编码器MAViE,通过融合多层特征、进行令牌路由及引入相关策略,在统一框架下实验,减少视觉令牌数量,提升多任务分数并降低处理时间,还给出模型设计与评估协议。

AI 中文摘要

视觉语言模型通常将预训练视觉编码器产生的所有令牌投影到大型语言模型中。然而,最终层特征可能会丢弃文本、局部属性和空间关系,而高分辨率输入会大幅增加上下文长度和推理延迟。我们引入了MAViE,一种多尺度自适应视觉编码器。MAViE使用位置相关门融合视觉Transformer的浅层、中层和深层特征,保留全局语义同时增强边缘、文本和局部结构。然后根据问题相关性、局部信息内容、全局语义和空间覆盖进行问题条件令牌路由,令牌预算适应图像复杂度。为减轻压缩损失,还引入全到压缩表示蒸馏和空间多样性正则化。在统一的7B语言模型框架下的模拟中,MAViE将SigLIP - SO400M视觉令牌平均数量从729减少到146(约80.0%),提高了VQAv2、GQA、TextVQA、ScienceQA - IMG和MMBench的平均分数2.2个百分点,同时将单图像到首个令牌时间从228毫秒减少到129毫秒。我们提供了完整模型设计和评估协议。

英文摘要

Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latency. We introduce \method, a Multi-scale Adaptive Vision Encoder. \method uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure. It then performs question-conditioned token routing according to question relevance, local information content, global semantics, and spatial coverage, with a token budget that adapts to image complexity. To mitigate compression loss, we further introduce full-to-compressed representation distillation and a spatial diversity regularizer. In an illustrative simulation under a unified 7B language-model framework, \method reduces the average number of SigLIP-SO400M visual tokens from 729 to 146 (approximately 80.0\%) and improves the mean score on VQAv2, GQA, TextVQA, ScienceQA-IMG, and MMBench by 2.2 percentage points, while reducing single-image time to first token from 228\,ms to 129\,ms. We provide the full model design and evaluation protocol. All reported numbers currently serve only as placeholders for paper organization and experimental design; formal claims require real training runs, independent replications, and official benchmark evaluation.

Comments10 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑