arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

层混合:用于视觉推理的动态层路由

Mixture of Layers: Dynamic Layer Routing for Visual Reasoning

Jeonghwan Kim, Sofia Stoica, Jiwan Chung, Ansel Blume, Hyeonjeong Ha, Zhenhailong Wang, Xin Luna Dong, Heng Ji

arXiv 2610.09440首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Yonsei University; Meta Reality Labs(伊利诺伊大学厄巴纳-香槟分校; 延世大学; Meta Reality Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对MLLMs视觉抽象与查询无关的问题,提出层混合(MoL)方法,通过指令条件动态路由聚合中间层特征,在7个细粒度任务上显著提升性能,实现查询自适应的视觉推理。

AI 中文摘要

预训练的视觉编码器包含逐层的视觉表示,这些表示在空间粒度、语义抽象程度以及对局部细节的敏感性方面各不相同。然而,大多数多模态大语言模型(MLLMs)仅依赖视觉编码器的最后一层或倒数第二层的表示,或采用固定的聚合规则,这使得视觉抽象在很大程度上与查询无关,并限制了对细粒度线索(如小物体、空间细节、文本和细微视觉属性)的访问。在这项工作中,我们提出了层混合(MoL),一种在视觉补丁级别上的指令条件层路由方法,它动态地聚合来自视觉编码器中间层的与查询相关的潜在表示。给定一个文本查询,MoL预测视觉编码器各层的路由概率,并在图像级别、补丁级别或通过混合路由机制对选定的隐藏状态进行top-k稀疏聚合。通过这种方式,MoL实现了对层特定视觉特征的查询自适应访问,以支持细粒度视觉推理。我们在7个细粒度视觉推理任务上的实验表明,性能显著提升,尤其是在细粒度视觉定位和理解任务上,与基线MLLMs相比,V*的整体准确率提高了+18.9%,HRBench4K提高了+4.5%,CharXiv提高了+16.3%,且无需采用多分辨率输入、简单交错多个视觉编码器或增加补丁令牌数量。我们研究了视觉编码器在不同层的感受野尺度及其采样行为,深入分析了为什么逐层采样是有帮助的,表明条件视觉表示是提升MLLMs视觉感知和推理能力的关键一步。我们的项目页面可在此https URL获取。

英文摘要

Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach at the visual patch level that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and performs a top-k sparse aggregation over selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. In doing so, MoL enables query-adaptive access to layer-specific visual features for fine-grained visual reasoning. Our experiments across 7 fine-grained visual reasoning tasks demonstrate substantial performance improvements, especially across fine-grained visual grounding and understanding tasks such as +18.9% improvement on V* in overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to the baseline MLLMs, without resorting to multi-resolution inputs, simple interleaving of multiple vision encoders, or increasing the number of patch tokens. We study vision encoders' receptive field scales across different layers and their sampling behaviors to provide an in-depth analysis of why layer-wise sampling is helpful, demonstrating that conditional visual representations are a key step towards better visual perception and reasoning in MLLMs. Our project page is available at https://wjdghks950.github.io/mol.github.io/.

CommentsNeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑